Speaking

Real incidents. Real lessons.
No vendor pitch.

I speak to engineering teams about how to run incidents with calm, clarity, and rigor. And about what AI actually changes in operations, which is less than the vendors claim and more than most teams have tried. Talks are drawn from 15+ years on the front line at eBay-scale e-commerce and fintech production, plus a live fleet where the AI operations layer is something I built and get paged for.

Topics

Incident Command from Startup to Enterprise

How the IC role scales (and breaks) as your engineering org grows. Practical frameworks for adopting ICS principles without enterprise overhead.

Incident Management That Actually Works

Severity triage, communication frameworks, escalation decision trees, and the runbook patterns that hold up at 3am.

Blameless Postmortems & Organizational Learning

How to run RCAs that produce durable corrective actions instead of theater, and how to keep them blameless when the pressure is on.

Site Reliability Engineering in Practice

SLOs, error budgets, toil elimination, and the reliability culture that separates teams that ship from teams that firefight.

Building the On-Call You'd Actually Want

Rotation design, alert quality, escalation paths, and how to make on-call a respected craft instead of an attrition driver.

The AI SRE: Killing Toil Without Handing Over the Keys

What actually works when you point AI at production (automated alert triage, self-healing remediation, agent-assisted incident response) and the structural controls that make it safe. Drawn from a live fleet where a model may rank a security response but can never name a target.

Postmortem Actions That Cannot Quietly Regress

Most corrective actions are a line in a document, which is the same as no control at all. How to turn a postmortem lesson into a continuously asserted invariant that pages the moment it stops being true, and how to tell which lessons earn one.

AI in Operations Without Giving Up Your Data

How AI is reshaping incident response, runbooks, and on-call, and why data sovereignty is the dividing line. Options range from private model deployments with providers like Google and Anthropic to fully self-hosted, internally trained models, and the tradeoffs that come with each.

Who I speak to

Enterprise platform teams

Maturing IC practice across multiple business units, tooling consolidation, and exec reporting.

Scale-ups & high-growth startups

First IC playbook, on-call rotation design, and the postmortem culture you set early.

Conferences & meetups

Keynotes and technical talks on incident response, SRE, and operational excellence.

Have an event in mind?

Send a note with the audience, format, and date. I'll come back with a topic shaped to fit.

jeremy@jeremymartinez.com