False Alarms Train You to Ignore Real Fires
Picture a fire crew that gets called to the same high school three times a day because teenagers keep pulling the alarm. The first few calls, they run. By the second week, they finish their coffee first. Most on-call rotations train their engineers the same way. Every alert that fires without needing a human teaches the person holding the pager that alerts don’t mean much.
What a false alarm costs
A false page costs the obvious things. Someone stops what they’re doing, opens a laptop, loads three dashboards, and finds nothing wrong. At 2 a.m., they lose the rest of the night’s sleep along with the twenty minutes. At 2 p.m., they lose the hour of focus they’d built up.
The bigger cost is what it teaches. After enough pages that resolve themselves, the on-call engineer builds a reflex: acknowledge, glance, wait for it to clear. Given the evidence, that reflex is reasonable. It’s also exactly the wrong reflex on the night the alert is real. The crew that strolled to the last forty false alarms strolls to the forty-first, and that one has smoke.
Flaky tests do the same thing to a CI pipeline. A suite with a handful of known-flaky tests trains the whole team to re-run instead of read, and the real failure gets lost in the noise. A noisy pager does the same damage, and it also wakes you up.
Actionable means someone has to act now
The test for a page is one question: when this fires, what does the on-call engineer do?
If the honest answer is “look at it and wait,” it shouldn’t page. If the answer is “nothing, it recovers on its own,” it shouldn’t page. If the answer is “file a ticket for Monday,” then the alert should file the ticket for Monday by itself, without waking anybody.
Some common offenders:
- CPU above 80% on one host in an autoscaling group. The autoscaler is already handling it.
- Disk at 70% on a volume that grows a gigabyte a month. That’s a ticket, maybe a quarter out.
- A single 500 error out of millions of requests.
- A batch job that runs late but still finishes before anyone needs the output.
Each of these is true and worth knowing. None of them needs a human awake right now. Google’s SRE book draws the same line: a page should be urgent, actionable, and about something users are feeling or are about to feel.
Page on what users feel
The alerts that hold up best watch symptoms. Users don’t care that CPU is high. They care that checkout is slow or failing. Alert on that, and let the causes show up on the dashboard you open after the page.
A symptom alert is usually a ratio with a duration attached. In Prometheus, for example:
yaml
- alert: CheckoutErrorRateHigh
expr: |
sum(rate(http_requests_total{service="checkout", code=~"5.."}[5m]))
/
sum(rate(http_requests_total{service="checkout"}[5m])) > 0.02
for: 10m
labels:
severity: page
annotations:
runbook: https://runbooks.example.com/checkout-errors
The ratio ignores the single stray 500. The for: 10m ignores the blip that clears on its own. The runbook link answers “what do I do?” before anyone has to ask it. If you can’t write the runbook for an alert, nobody knows what action it’s asking for, and it shouldn’t be paging anyone.
Prune the pager like code
Alert rules pile up. Someone adds one after an incident, the incident gets fixed, and the alert stays forever. Nobody deletes alerts, because deleting one feels like the opening scene of the next outage.
Treat the alert config like any other code with a maintenance cost. Every alert gets an owner and a runbook. Every page gets reviewed at the end of the rotation with one question: did someone take action? If nobody did, the alert gets fixed, demoted to a ticket, or deleted.
Split severity in two. Pages are for things that need a human now. Tickets are for things that need a human eventually. Most of what wakes people up belongs in the second bucket.
If you acknowledged a page and went back to sleep, it should never have woken you.
Keep the crew running
The fire department’s answer to the school is to fix the alarm: put a cover on the pull station and find out who keeps pulling it. The crew keeps running because the alarm still means something.
Your pager works the same way. Every alert you delete or demote makes the ones that remain worth running for. Cut the false alarms, and the next time the pager goes off, people will run.