Work
Systems I designed, built, and still operate.
Not case studies written after the fact. These are live systems carrying real traffic and regulated data, and between them they close every phase of an incident: detect, diagnose, respond, govern who is allowed to act, recover, and learn in a way that sticks. Each deep dive carries the decisions I made, the ones I got wrong first, and the artifact that proves the claim rather than a paragraph asking you to take my word for it.
Round-the-clock operations is an arithmetic problem before it is a technology one. One seat covered continuously is 4.2 people at forty hours a week, and more once holidays, sickness and attrition are real. The last time I staffed that properly it was twelve operators on a 24x7 desk. This fleet has no rotation at all, and the first action on an incident still happens before anyone would have woken up.
Written at capability level: architecture, reasoning and measured numbers, with no hostnames, addresses or paths. The same rule the systems themselves enforce.
What it replaces
| One seat, covered around the clock | 4.2 people |
|---|---|
| The NOC I ran at New Frontier | 12 people |
| This fleet | 1 people |
| Removed by automation | 90 percent |
|---|---|
| Still done by hand | 10 percent |
Response time is the product
Every minute before the first action is a minute somebody is deciding whether to trust you again. The largest term in that number was never diagnosis. It was finding a human, waking them, and waiting for context they had to go collect.
Find it before the client does
The worst version of this job is learning your service is down from the person paying for it. Checks here assert the thing works rather than that a process is listening, because a green dashboard above an unhappy customer is the most expensive lie in operations.
Toil is what it takes from you
Toil is not merely inefficient. It is the reason good engineers leave, and it eats the hours that would have gone into the work that moves something forward. Take it off the rotation and you get those hours back. That is the whole pitch.
Security & Response
The AI SOC
Everybody has detection. Nobody has anybody to read it. That is not a problem you hire your way out of, and it is not a problem you buy another product to fix.
- One question per case: is it real, and by what mechanism. Everything else is decided elsewhere.
- The investigator holds no tools and gathers no evidence itself. It asks for evidence by number.
- It can close a case. It can never raise a severity.
Autonomous Response
The AI never writes a command. It picks a number off a list, so the worst a hostile input can do is make it pick a different approved action.
- No field in the protocol can carry a host, a path or a command. A target is not something the validator catches, it is something the schema cannot express.
- Ten gates stand between a chosen option and a state change, and any of them can stop it.
- Destructive actions are human only. Tested by trying, not assumed.
NDR Rollout
The value was never inbound attack detection. It was being able to prove whether data left.
- The vendor's stock severity mapping could never have paged anyone. Found by testing it before installing anything.
- Four staged deployments, and the one host left out is a written decision rather than an omission.
- Six verification tests, each designed so that it could fail. One did, immediately.
Reliability & Automation
The Fleet Control Plane
One person cannot watch a production fleet around the clock. So the machine takes the first action, and the interesting part is everything that stops it.
- Covering one seat around the clock takes four to five people. This runs on one, and the first action does not wait for me to wake up.
- Every postmortem finding ships as a check that re-proves itself on a schedule.
- No change surprises the on-call: work is declared, and suppression is one host for one window.
The Self-Healing NOC & War Room
A 200 means something is listening. It does not mean the site is up. That gap is where an outage sits quietly while every dashboard stays green.
- A staffed first-response rotation for nights and weekends is four to five people. This is none of them, and it answers in seconds.
- A recovery counts only when the service serves its own correct content, never on a status code.
- One safe attempt, then it escalates with the evidence already gathered.
Backups Built For The Day You Are Already Compromised
Everyone has backups. The question worth asking is whether the thing that just encrypted your fleet can also reach them, and for most people the honest answer is yes.
- Three copies, two media, one offsite, and the offsite copy is pulled by a machine the fleet cannot reach.
- No host on the fleet holds a credential that can delete a backup.
- Restores are drilled daily. The number I keep is when one last succeeded.
The Fleet Worker Pool
Any machine can do any job, a stranger's code runs somewhere with no home to go back to, and the list of what a machine can do is not allowed to lie about it.
- One worker binary runs every project's jobs, so there is one attack surface to hold instead of one per project.
- Sandboxed code reaches no secret, because the secrets are not in its namespace to deny.
- Routing heals itself across sixty models as providers throttle, retire and restore them.
The Terminal Brain
A committee of models reads over your shoulder in a live production shell. The watching tier cannot type. Not switched off, not configured down. There is no code path that types.
- Three independent lenses read the screen, and disagreement between them forces a human decision.
- The watching tier cannot type. Not switched off, not configured down. There is no code path that types.
- One keystroke drops it to watch-only. No confirmation, no grace period, no setting to disable it.
Platforms & Products
HitDirector & WebHostNOC
Most hosting sells you disk space and wishes you luck. This sells the part nobody advertises, which is somebody answering at two in the morning.
- Round-the-clock incident response on a fleet with no night shift, because the first response is automated rather than staffed.
- Backups verified by restore drill, daily, and a drill that stops running goes red.
- No secret in any repository. Sealed storage or an on-host file, nothing else.
The CPA Client Portal
A stolen copy of the database is a pile of ciphertext. The keys that open it are on a different machine, and that machine does not run the website.
- One firm per machine, so the isolation is physical and an auditor can see the boundary.
- No decryption key on the app host, so a stolen database is a pile of ciphertext.
- One idempotent command stands up a tenant, which is what makes it a product rather than six bespoke builds.
The White-Label Telehealth Portal
One codebase, many clinics, and a rule that a tenant may differ in its branding and in nothing else.
- One core codebase. A clinic differs in branding and in nothing else, so a security fix reaches all of them in one deploy.
- No patient data on the media host. Compromising the relay gets you a busy relay and nothing to read.
- One clinic per machine, so a bad day is one clinic's bad day.
SignFlow: The E-Signature Service
Renting your compliance story from a signature vendor means the evidence lives in their account, under their breach policy, on their retention schedule.
- A signature is an evidentiary claim needing four things: consent, intent, attribution, and the exact document version.
- No third party holds the evidence, so it is not on somebody else's retention schedule or breach policy.
- Every event is hash-chained, so altering history means rewriting all of it, and that is visible.
Incident Commander HQ
The training I wish had existed the first time I was handed a global outage.
- Fifteen years of real incidents behind the curriculum, not a framework summarized from a book.
- Free field guide, and no vendor pitch. Training that becomes a funnel starts teaching the tool.
Want the walkthrough instead of the write-up?
I'll screen-share the live consoles, the SOC board, the war room, the invariant assertions, the terminal committee, and answer questions against running systems.
Request a walkthrough