Work

Systems I designed, built, and still operate.

Not case studies written after the fact. These are live systems carrying real traffic and regulated data, and between them they close every phase of an incident: detect, diagnose, respond, govern who is allowed to act, recover, and learn in a way that sticks. Each deep dive carries the decisions I made, the ones I got wrong first, and the artifact that proves the claim rather than a paragraph asking you to take my word for it.

Round-the-clock operations is an arithmetic problem before it is a technology one. One seat covered continuously is 4.2 people at forty hours a week, and more once holidays, sickness and attrition are real. The last time I staffed that properly it was twelve operators on a 24x7 desk. This fleet has no rotation at all, and the first action on an incident still happens before anyone would have woken up.

Written at capability level: architecture, reasoning and measured numbers, with no hostnames, addresses or paths. The same rule the systems themselves enforce.

What it replaces

People required for continuous coverageCovering one seat around the clock takes 4.2 people at forty hours a week. The network operations centre took twelve. This fleet runs on one.One seat, covered around the clockOne seat, covered around the clock: 4.2 people4.2The NOC I ran at New FrontierThe NOC I ran at New Frontier: 12 people12This fleetThis fleet: 1 people1
People required for continuous coverage
One seat, covered around the clock4.2 people
The NOC I ran at New Frontier12 people
This fleet1 people
Nobody puts this on a slide, so here it is. Keeping one seat staffed around the clock is 4.2 people before anyone takes a holiday, gets sick, or quits in March. I ran that desk with twelve operators and I know exactly what it costs, in money and in the people doing it. This fleet has none of that, and the first action on an incident still lands before a phone would have finished ringing.
Manual work removed from a twelve-person team at eBayAutomation removed ninety percent of the manual operational work carried by a twelve-person team, so the pages that still fired were genuinely novel.Manual work on a twelve-person teamTotal manual work carried by the team90 percent removed by automation90% removedWhat was left was the work that actually needed a person.
Manual work on a twelve-person team at eBay
Removed by automation90 percent
Still done by hand10 percent
This is the older half of the argument, and it is not new. At eBay I automated 90% of the manual work off a twelve-person team, years before any of this had a fashionable name. What changed is not the ambition. It is that the machine can now read an alert, form a view, and act inside a boundary I drew, which is the part that used to require a person awake at four in the morning.

Response time is the product

Every minute before the first action is a minute somebody is deciding whether to trust you again. The largest term in that number was never diagnosis. It was finding a human, waking them, and waiting for context they had to go collect.

Find it before the client does

The worst version of this job is learning your service is down from the person paying for it. Checks here assert the thing works rather than that a process is listening, because a green dashboard above an unhappy customer is the most expensive lie in operations.

Toil is what it takes from you

Toil is not merely inefficient. It is the reason good engineers leave, and it eats the hours that would have gone into the work that moves something forward. Take it off the rotation and you get those hours back. That is the whole pitch.

Security & Response

Reliability & Automation

The Fleet Control Plane

One person cannot watch a production fleet around the clock. So the machine takes the first action, and the interesting part is everything that stops it.

  • Covering one seat around the clock takes four to five people. This runs on one, and the first action does not wait for me to wake up.
  • Every postmortem finding ships as a check that re-proves itself on a schedule.
  • No change surprises the on-call: work is declared, and suppression is one host for one window.

The Self-Healing NOC & War Room

A 200 means something is listening. It does not mean the site is up. That gap is where an outage sits quietly while every dashboard stays green.

  • A staffed first-response rotation for nights and weekends is four to five people. This is none of them, and it answers in seconds.
  • A recovery counts only when the service serves its own correct content, never on a status code.
  • One safe attempt, then it escalates with the evidence already gathered.

Backups Built For The Day You Are Already Compromised

Everyone has backups. The question worth asking is whether the thing that just encrypted your fleet can also reach them, and for most people the honest answer is yes.

  • Three copies, two media, one offsite, and the offsite copy is pulled by a machine the fleet cannot reach.
  • No host on the fleet holds a credential that can delete a backup.
  • Restores are drilled daily. The number I keep is when one last succeeded.

The Fleet Worker Pool

Any machine can do any job, a stranger's code runs somewhere with no home to go back to, and the list of what a machine can do is not allowed to lie about it.

  • One worker binary runs every project's jobs, so there is one attack surface to hold instead of one per project.
  • Sandboxed code reaches no secret, because the secrets are not in its namespace to deny.
  • Routing heals itself across sixty models as providers throttle, retire and restore them.

The Terminal Brain

A committee of models reads over your shoulder in a live production shell. The watching tier cannot type. Not switched off, not configured down. There is no code path that types.

  • Three independent lenses read the screen, and disagreement between them forces a human decision.
  • The watching tier cannot type. Not switched off, not configured down. There is no code path that types.
  • One keystroke drops it to watch-only. No confirmation, no grace period, no setting to disable it.

Platforms & Products

HitDirector & WebHostNOC

Most hosting sells you disk space and wishes you luck. This sells the part nobody advertises, which is somebody answering at two in the morning.

  • Round-the-clock incident response on a fleet with no night shift, because the first response is automated rather than staffed.
  • Backups verified by restore drill, daily, and a drill that stops running goes red.
  • No secret in any repository. Sealed storage or an on-host file, nothing else.

The CPA Client Portal

A stolen copy of the database is a pile of ciphertext. The keys that open it are on a different machine, and that machine does not run the website.

  • One firm per machine, so the isolation is physical and an auditor can see the boundary.
  • No decryption key on the app host, so a stolen database is a pile of ciphertext.
  • One idempotent command stands up a tenant, which is what makes it a product rather than six bespoke builds.

The White-Label Telehealth Portal

One codebase, many clinics, and a rule that a tenant may differ in its branding and in nothing else.

  • One core codebase. A clinic differs in branding and in nothing else, so a security fix reaches all of them in one deploy.
  • No patient data on the media host. Compromising the relay gets you a busy relay and nothing to read.
  • One clinic per machine, so a bad day is one clinic's bad day.

SignFlow: The E-Signature Service

Renting your compliance story from a signature vendor means the evidence lives in their account, under their breach policy, on their retention schedule.

  • A signature is an evidentiary claim needing four things: consent, intent, attribution, and the exact document version.
  • No third party holds the evidence, so it is not on somebody else's retention schedule or breach policy.
  • Every event is hash-chained, so altering history means rewriting all of it, and that is visible.

Incident Commander HQ

The training I wish had existed the first time I was handed a global outage.

  • Fifteen years of real incidents behind the curriculum, not a framework summarized from a book.
  • Free field guide, and no vendor pitch. Training that becomes a funnel starts teaching the tool.

Want the walkthrough instead of the write-up?

I'll screen-share the live consoles, the SOC board, the war room, the invariant assertions, the terminal committee, and answer questions against running systems.

Request a walkthrough