All work

Platforms & Products

Incident Commander HQ

The training I wish had existed the first time I was handed a global outage.

Built and maintainNext.jsContent platformDockerCloudflare
  • Fifteen years of real incidents behind the curriculum, not a framework summarized from a book.
  • Free field guide, and no vendor pitch. Training that becomes a funnel starts teaching the tool.

The problem

Most engineering organizations are still improvising incident command the first time the pager goes off at 3am, and the cost of that improvisation is paid in minutes of downtime and years of team attrition.

Incident command is a learnable skill with a fifty-year body of practice behind it. It just was not written down for software.

What it covers

ICS principles translated to production engineering, severity triage, escalation decision trees, stakeholder and executive communication during an active incident, the runbook patterns that hold up under pressure, and blameless postmortem workflows that produce durable corrective actions instead of theater.

Engineering decisions

The calls I made, and what each one cost.

No vendor pitch.

The moment incident training becomes a funnel for a tool, it starts teaching the tool instead of the discipline.

More systems

I'm looking for Incident Commander and SRE roles.

If your team is drowning in toil, alert noise, or incidents that never quite close. That is the work I do.

Get in touch