Matthew Boston

Leave On-Call Better Than You Found It

August 26, 2024

Keeping production up is the obvious part of being on call. The part that pays off over a year is making the rotation suck less for whoever carries the pager next. On any rotation, that person is eventually you.

The next person is you

The math is simple. On a six-person weekly rotation, you’re back in six weeks. Every annoyance you shrug off this week is waiting for you then: the alert with no runbook, the deploy that needs a manual cache flush, the dashboard nobody can find.

Most on-call pain repeats. The week you’re on call is the week you have both the motivation and the context to fix it. Next week you’ll be back on feature work, and the fix will slide until you’re paged for the same thing again.

Write it down while it’s fresh

At 3 a.m. you’ll figure something out: the log query that shows the failing requests, the service that has to restart before the other one, the dashboard with the useful graph. By 10 a.m. you’ll remember the shape of it and forget the query.

Write it down during the incident, in the channel, as you go. Then turn it into a runbook before the rotation ends. A good runbook entry is boring and specific. Something like this:

```markdown ## Payment worker queue backing up

  1. Check queue depth: bin/queue-stats payments
  2. If depth > 10k and workers are healthy, scale up: kubectl scale deploy/payment-worker --replicas=12
  3. If workers are crash-looping, check the last deploy and roll back: bin/rollback payment-worker ```

Exact commands, exact thresholds. “Investigate the queue” is a wish. The person reading the runbook is half-awake and may have never seen this failure before, so write for that person.

Script it the second time

The first time you do something by hand on call, write it down. The second time, script it. Once the script does the same thing every run, ask whether a human needs to run it at all.

The runbook above has a step that’s a perfect candidate. “If depth > 10k and workers are healthy, scale up” is a rule a machine can evaluate. Turn it into an autoscaling policy and that page stops happening. Every piece of on-call toil you automate raises the floor for everyone on the rotation, including the engineer who joined last month and has never seen this service fail.

Fix what woke you up

Every page you get is feedback on the alerting. At the end of the rotation, go through the list and give each page a verdict:

  • Real and actionable: make sure the runbook covers it.
  • Real, but it could have waited until morning: demote it to a ticket.
  • Noise: fix the threshold or delete the alert.

The pages nobody acted on are the ones that make on-call miserable. They’re also the easiest to fix, because the person who just got woken up by them has all the context. In two weeks, nobody will remember why that alert fired four times on Tuesday night.

Make it part of the job

None of this happens if on-call improvements compete with sprint work and lose every time. A few habits help:

  • The on-call engineer doesn’t carry sprint work. Their project for the week is on-call itself: interrupts first, improvements in between.
  • Each rotation ends with a short handoff note covering what paged, what got fixed, and what’s still open.
  • The team keeps a standing on-call backlog that each rotation pulls from, so improvements don’t depend on someone remembering.

The handoff note is the cheapest of these and does the most. It takes ten minutes and hands the next person everything you learned that week.

Six weeks from now

The rotation you’re on today is the one previous on-call engineers built, one fix at a time, or failed to build. You won’t fix all of it in a week. Fix one thing, write down two more, and hand the pager over a little lighter than it was handed to you.

In six weeks, the pager comes back to you.