Resume

Jeremy Martinez

Senior Site Reliability Engineer · Incident Commander · AI Operations

Las Vegas, NV · (720) 310-5673 · jeremy@jeremymartinez.com

Download PDF

Professional summary

Senior SRE and Incident Commander, 15+ years on mission-critical production systems, including 11 at eBay holding a 99.997% uptime SLA and taking 90% of the manual toil off a 12-person team. I have spent that career driving toil and mean time to recovery down, and AI is the first tool that moves both at once: coverage that would take four to five people to staff continuously runs unattended on my own fleet, and the first action on an incident does not wait for anyone to wake up. Today I operate a multi-client cloud estate spanning consumer, healthcare and public-sector workloads, and prove the automation on that fleet before it reaches a client.

Core skills

Incident Response & Observability
Incident command (ICS), major incident management, postmortems and RCA, MTTR reduction, 24x7 on-call, game days, Datadog, Prometheus, Grafana, SysDig, Sumo Logic, PagerDuty, Rootly, ServiceNow, Jira
AI & Agentic Operations
AI incident response, LLM alert triage, autonomous remediation, agent design with constrained action spaces, multi-model quorum voting, prompt-injection resistance by construction (OWASP LLM Top 10), sandboxed tool execution, human-in-the-loop approval gates, replay-based regression testing
Security Operations & Compliance
Wazuh (SIEM/HIDS), SOAR, Suricata NDR, auditd, Falco/eBPF, Cloudflare WAF, fail2ban, secrets management (OpenBao/Vault), GLBA FTC Safeguards, IRS Pub 4557 WISP, ESIGN/UETA
Cloud, Containers & IaC
AWS, Azure, GCP, Kubernetes, OpenShift, Terraform, CloudFormation, Helm, Ansible, Jenkins, Argo CD, CI/CD
Reliability & Infrastructure
High availability, disaster recovery, capacity planning, load balancing, traffic engineering, DDoS mitigation, Linux (RHEL, CentOS, Ubuntu), Solaris, MySQL, PostgreSQL, Ceph, NFS, iSCSI, Veritas VCS
Languages
Python, Bash, Shell, Perl, PHP, Java

Selected Engineering: AI Operations & Incident Response

Systems I designed, built, and operate on a production fleet of my own. Proven against the strictest regimes I work under, GLBA Safeguards, IRS Pub 4557 and HIPAA, because an outage does not care what industry the data belongs to. Full architecture write-ups at jeremymartinez.com/work.

Detection and alerting

Raw security telemetry becomes a short working set of attacker-scoped cases with the evidence already attached, so triage starts at understanding rather than at collection.

Automated incident response

Continuous coverage without a rotation. Containment starts before anyone is paged, which removes the largest term in mean time to recovery. The model returns an identifier for one pre-authorized action, and no field in the schema can carry a host or a command, so an injected response is something the protocol cannot express.

Authority and approval controls

Any voter can lower an outcome and none can raise it, approval carries a second factor verified on the privileged side, and a request that times out does nothing rather than proceeding.

Recovery and post-incident follow-through

Backups are append-only and pulled offsite by a machine the fleet holds no credential to reach. Every postmortem ships as an executable invariant re-proved on a schedule, which is continuous control monitoring rather than an annual assertion.

Products built on the same fleet

Website Checkup renders a site in a real browser and grades nine dimensions including WCAG accessibility, driven by an unattended worker. A local events guide reuses that worker to verify listings against each venue's own site before anything publishes. Same agent design, unrelated problems, which is the argument that it generalizes.

Incident Commander Field Guide & Simulator · incidentcommandhq.com

Public incident-command training platform and live-scenario simulator. ICS roles, severity triage, stakeholder communication, blameless postmortems.

Experience

Dynascale Inc.

03/2024 - Present

Senior Site Reliability Engineer, Incident Commander

  • Incident Commander for customer and platform incidents across a multi-client AWS, Azure, and GCP estate spanning consumer, healthcare and public-sector workloads. Triage, mitigation, stakeholder escalation, and post-incident remediation end to end.
  • Cut manual operational intervention 30% with self-healing workflows, custom alerting pipelines, real-time telemetry, and infrastructure lifecycle automation in Terraform, CloudFormation, and Ansible.
  • Reduced cloud spend through reserved-instance strategy, autoscaling tuning, rightsizing, and provisioning changes across three clouds.
  • Bringing LLM-assisted remediation to recurring system-administration incidents across Hyper-V, AWS, and Azure. The bounded-action and human-approval patterns proven on my own regulated fleet.
  • Own disaster recovery (backup validation, failover testing, and response playbooks) and converted undocumented tribal knowledge into standard SOPs and runbooks.
  • Engineered automated patch and upgrade workflows across cloud and on-premises fleets, eliminating manual patching; mentor engineers on reliability practice and production ownership.

Upstart Inc.

07/2022 - 02/2024

Senior Site Reliability Engineer, Incident Commander

  • Incident Commander for enterprise production incidents at a public consumer-lending platform. Ran the bridge, assigned roles, drove mitigation, and owned executive stakeholder communication through major outages on a 24x7 rotation.
  • Owned Rootly end to end (configuration, incident lifecycle workflows, escalation automation) and standardized severity definitions and response process across the organization, driving adoption with engineering teams that did not report to me, which is the only kind of authority an Incident Commander has.
  • Led blameless postmortems and formal root cause analyses, converting findings into durable corrective actions rather than one-off fixes.
  • Delivered weekly reliability metrics and incident analytics to executive leadership.
  • Designed and ran incident simulation exercises and gamedays, and maintained the runbook and playbook library the on-call organization responded from.

eBay Inc.

10/2011 - 07/2022

Production Unix Systems Engineer / MTS / Incident Responder

  • Held a 99.997% uptime SLA across critical services as senior Incident Responder and escalation owner for large-scale production incidents on a global e-commerce platform processing tens of thousands of dollars per second at peak.
  • Designed automation that took 90% of the manual operational toil off a 12-person team. Recognized with a Critical Talent Bonus for incident management and automation impact.
  • Led real-time triage, coordination, mitigation, and root cause analysis spanning infrastructure, application, database, and network domains.
  • Built automated monitoring, alerting, and self-healing workflows that resolved disk, CPU, service, and capacity failures before they paged.
  • Operated large-scale Unix infrastructure (RHEL, CentOS, Solaris, Ubuntu), Veritas VCS clusters fronting high-availability Oracle, Akamai CDN integrations, and a 10,000+ node Hadoop cluster.

New Frontier Media Inc.

08/2008 - 10/2011

Systems Engineer

  • Built and ran the network operations center: twelve operators on 24x7 monitoring and first response. Owned shift scheduling, coverage, technical direction, and the training that brought new operators up to the standard.
  • Designed and operated high-traffic streaming platforms supporting national broadcast distribution.
  • Built and operated clustered outbound mail infrastructure at scale: SMTP server clusters, sender authentication with SPF and domain keys, IP rotation and domain management for deliverability and sender reputation.
  • Built and scaled infrastructure for social and e-commerce platforms on MongoDB, Redis, Nginx, Node.js, VMware, and Xen, with clustered MySQL.
  • Led web application security assessments and incident response.

Hit Director Domains · Comtech · ISSG · The Internet Web Hosting Co.

1997 - 2008

Systems / Network Engineering (condensed)

  • Built and operated high-traffic Linux server farms, web hosting platforms, and secure streaming systems, including PCI-compliant and hardened production environments.
  • Administered enterprise Unix, load balancers, and automation frameworks; led migrations, performance tuning, and vulnerability remediation.

US Army

1993 - 1997

Communications Center Operator

  • Operated secure telecommunications, satellite and fiber systems, and cryptographic equipment in TS/SCI environments; multiple Army Achievement Medals.

Depth on request

Nine of these bullets have a full write-up behind them.

Architecture, the engineering decisions and the reasoning behind them, the verification that proved each one, and the things I got wrong the first time. Written at capability level, with no hostnames, addresses or paths.

Read the deep dives

Education

Management / Computer Information Systems
Park University, Parkville, Missouri

Certifications

  • PagerDuty Incident Response (2024)
  • Datadog Monitoring AWS (2023)
  • Sumo Logic Fundamentals and Search Mastery (2022)
  • CompTIA Security+