Let an AI Agent Take the First Page
Most on-call rotations spend skilled engineers on Level 1 work: the same incidents, fixed with the same runbooks, every time. Make an autonomous AI agent your Level 1. Give it the logs, the systems, the code, and the runbooks. Let it resolve what it can and page a human only when it’s stuck. The human job changes from pager zombie to system teacher.
Pager zombies
Pull a quarter of pages for most services and the same handful of incidents show up again and again. A disk filling with logs. A worker stuck on a poisoned message. A queue backing up after a traffic spike. A deploy that raised the error rate and needs rolling back. Each one has a runbook, and the runbook has steps.
The engineer who gets paged for these can do much harder work. At 3 a.m. that doesn’t matter. They’re a pager zombie: half-awake, following a checklist, pasting commands out of a wiki page. Then they lose the rest of the night and spend the next day at half speed.
Following a runbook is a poor use of a senior engineer and a good use of an agent.
What Level 1 needs
An agent working an incident needs what a human needs, through tools instead of browser tabs:
- Logs, metrics, and traces it can query directly.
- The deploy history and the code, so it can connect “error rate went up at 14:02” to “this commit shipped at 14:00.”
- The runbooks, written as instructions an agent can follow.
That last one is mostly a format change. A runbook is a SKILL.md written for humans: when this alert fires, check this, then do that. Rewrite it with exact commands and explicit decision points, and it works for both readers.
Guardrails before write access
Read access is easy to justify. Write access to production is where this gets serious, and the agent should earn it the same way you’d earn auto-merge for trivial PRs: write down exactly what’s allowed, put the rules in a file, and enforce them outside the agent.
Start with a short allowlist of remediations. An illustrative policy might look like this:
yaml
allowed_actions:
- restart_pod
- scale_deployment: { max_replicas: 20 }
- rollback_deploy: { within_minutes: 60 }
- toggle_feature_flag
- clear_disk: { paths: ["/var/log/app"] }
escalate_after_minutes: 15
never:
- database_migrations
- iam_changes
- deleting_data
Everything outside the list escalates. Every action gets posted to the incident channel as it happens, so there’s a record a human can read later. A time budget caps how long the agent works before a person hears about it, whether or not it thinks it’s close.
Enforce the allowlist in the tooling the agent calls, where the agent can’t talk its way past it. It’s the same principle as the harness around a coding agent. The agent never decides whether the guardrails apply. It only decides what to do inside them.
Escalation with a head start
When the agent gets stuck, the page it sends should be better than the one it received. A raw alert says CheckoutErrorRateHigh. An agent’s escalation can say what fired, what it checked, what it ruled out, what it tried, and what it thinks is going on.
The human who gets that page starts from a briefing instead of a blank dashboard. Even on the incidents the agent can’t resolve, it has already done the first fifteen minutes of investigation, and it did them while the human was still finding their glasses.
From pager zombie to system teacher
Every escalation exposes a gap in what the agent knows or is allowed to do. Maybe the runbook doesn’t exist. Maybe the log query it needed isn’t exposed as a tool. Maybe the fix was safe but wasn’t on the allowlist.
After an escalation, the human’s job is to close that gap. Write the runbook, add the tool, or widen the guardrail where the evidence says it’s safe. Each fix moves another class of incident from “escalate” to “resolved.” Over time, the share of pages a human sees goes down, and the ones that remain are the novel failures that need human judgment.
That changes what on-call asks of people. The work moves from executing checklists at night to improving the system during the day, which is the work senior engineers should have been doing all along.
The agent takes the 3 a.m. page about the full disk. You take the one nobody has seen before, with the briefing already written.