The loop, always in this order
| Staffed rotation | Detect, Page, Wake up, Open laptop, Gather context, First action |
|---|---|
| Unattended first response | Detect, First action, Verify, Escalate with evidence |
1. Verify the symptom yourself
Don't trust one signal. Re-probe. If it is already healthy, resolve it as a flap and take no action.
2. Gather a read-only evidence bundle
Host reachability, disk, container state and restart counts, metrics context, recent logs. Scrubbed, never echoing record contents.
3. Ask what changed
Last patch run, recent snapshots, time since deploy, and whether a declared maintenance window on this host just opened or closed. A blip inside an announced change is expected, not an incident.
4. Classify
Host unreachable, container down or crash-looping, resource exhaustion, TLS, application 5xx, or dependency failure.
5. Crash-loop guard
If it has already restarted many times recently, do not restart it again. That masks a real fault. Escalate.
6. Smallest safe reversible action
Snapshot first on any host that has snapshots. One attempt. Never loop-retry a failed action.
7. Re-verify by content, not status code
See below. This is the rule that took the longest to learn.