Detection and alerting
Raw security telemetry becomes a short working set of attacker-scoped cases with the evidence already attached, so triage starts at understanding rather than at collection.
Resume
Senior Site Reliability Engineer · Incident Commander · AI Operations
Las Vegas, NV · (720) 310-5673 · jeremy@jeremymartinez.com
Senior SRE and Incident Commander, 15+ years on mission-critical production systems, including 11 at eBay holding a 99.997% uptime SLA and taking 90% of the manual toil off a 12-person team. I have spent that career driving toil and mean time to recovery down, and AI is the first tool that moves both at once: coverage that would take four to five people to staff continuously runs unattended on my own fleet, and the first action on an incident does not wait for anyone to wake up. Today I operate a multi-client cloud estate spanning consumer, healthcare and public-sector workloads, and prove the automation on that fleet before it reaches a client.
Systems I designed, built, and operate on a production fleet of my own. Proven against the strictest regimes I work under, GLBA Safeguards, IRS Pub 4557 and HIPAA, because an outage does not care what industry the data belongs to. Full architecture write-ups at jeremymartinez.com/work.
Raw security telemetry becomes a short working set of attacker-scoped cases with the evidence already attached, so triage starts at understanding rather than at collection.
Continuous coverage without a rotation. Containment starts before anyone is paged, which removes the largest term in mean time to recovery. The model returns an identifier for one pre-authorized action, and no field in the schema can carry a host or a command, so an injected response is something the protocol cannot express.
Any voter can lower an outcome and none can raise it, approval carries a second factor verified on the privileged side, and a request that times out does nothing rather than proceeding.
Backups are append-only and pulled offsite by a machine the fleet holds no credential to reach. Every postmortem ships as an executable invariant re-proved on a schedule, which is continuous control monitoring rather than an annual assertion.
Website Checkup renders a site in a real browser and grades nine dimensions including WCAG accessibility, driven by an unattended worker. A local events guide reuses that worker to verify listings against each venue's own site before anything publishes. Same agent design, unrelated problems, which is the argument that it generalizes.
Public incident-command training platform and live-scenario simulator. ICS roles, severity triage, stakeholder communication, blameless postmortems.
Senior Site Reliability Engineer, Incident Commander
Senior Site Reliability Engineer, Incident Commander
Production Unix Systems Engineer / MTS / Incident Responder
Systems Engineer
Systems / Network Engineering (condensed)
Communications Center Operator
Depth on request
Architecture, the engineering decisions and the reasoning behind them, the verification that proved each one, and the things I got wrong the first time. Written at capability level, with no hostnames, addresses or paths.
Read the deep divesManagement / Computer Information Systems
Park University, Parkville, Missouri