Speaking
Real incidents. Real lessons.
No vendor pitch.
I speak to engineering teams about how to run incidents with calm, clarity, and rigor. And about what AI actually changes in operations, which is less than the vendors claim and more than most teams have tried. Talks are drawn from 15+ years on the front line at eBay-scale e-commerce and fintech production, plus a live fleet where the AI operations layer is something I built and get paged for.
Topics
Incident Command from Startup to Enterprise
How the IC role scales (and breaks) as your engineering org grows. Practical frameworks for adopting ICS principles without enterprise overhead.
Incident Management That Actually Works
Severity triage, communication frameworks, escalation decision trees, and the runbook patterns that hold up at 3am.
Blameless Postmortems & Organizational Learning
How to run RCAs that produce durable corrective actions instead of theater, and how to keep them blameless when the pressure is on.
Site Reliability Engineering in Practice
SLOs, error budgets, toil elimination, and the reliability culture that separates teams that ship from teams that firefight.
Building the On-Call You'd Actually Want
Rotation design, alert quality, escalation paths, and how to make on-call a respected craft instead of an attrition driver.
The AI SRE: Killing Toil Without Handing Over the Keys
What actually works when you point AI at production (automated alert triage, self-healing remediation, agent-assisted incident response) and the structural controls that make it safe. Drawn from a live fleet where a model may rank a security response but can never name a target.
Postmortem Actions That Cannot Quietly Regress
Most corrective actions are a line in a document, which is the same as no control at all. How to turn a postmortem lesson into a continuously asserted invariant that pages the moment it stops being true, and how to tell which lessons earn one.
AI in Operations Without Giving Up Your Data
How AI is reshaping incident response, runbooks, and on-call, and why data sovereignty is the dividing line. Options range from private model deployments with providers like Google and Anthropic to fully self-hosted, internally trained models, and the tradeoffs that come with each.
Who I speak to
Enterprise platform teams
Maturing IC practice across multiple business units, tooling consolidation, and exec reporting.
Scale-ups & high-growth startups
First IC playbook, on-call rotation design, and the postmortem culture you set early.
Conferences & meetups
Keynotes and technical talks on incident response, SRE, and operational excellence.
Have an event in mind?
Send a note with the audience, format, and date. I'll come back with a topic shaped to fit.
jeremy@jeremymartinez.com