Matthew Boston

Stop Counting Incidents

August 8, 2024

Reward a team for having fewer incidents and they will have fewer incidents, on paper. The outages still happen. They just get declared later, scoped smaller, or never declared at all. Measure how quickly a team resolves incidents instead, and the incentives point at the work that limits the damage.

What an incident count rewards

Goodhart’s law, in Marilyn Strathern’s phrasing: when a measure becomes a target, it ceases to be a good measure. Incident count is a textbook case.

Put “number of incidents” on a quarterly scorecard and watch what changes. Engineers stop declaring early. Someone sees error rates climbing and says “let’s give it ten minutes and see if it clears.” A partial outage becomes a “degradation.” A Sev 2 becomes a Sev 3 because Sev 3s don’t count. Nobody is lying, exactly. They’re playing to the scoreboard you gave them.

The minutes between “something looks wrong” and “we’ve declared an incident” are some of the most expensive in the whole timeline. Customers are already affected, and nobody is coordinating the response. An incident count makes that gap longer.

Declaring should be cheap

If you want people to declare, make declaring cost nothing. One chat command should open a channel, page the right rotation, and start a timeline:

text /incident declare sev3 "checkout latency elevated in us-east"

Then make it safe to be wrong. An incident that turns out to be nothing gets closed with a one-line note, and nobody asks why it was declared. A false declaration costs a few minutes. A late declaration costs whatever the outage costs while nobody is coordinating. Everyone on the team should understand that asymmetry, and the way to teach it is to never punish the cheap mistake.

Time to recover is the number that matters

The DORA research program has tracked time to restore service for years as one of its four key metrics, next to deployment frequency, lead time for changes, and change failure rate. Failures are going to happen. What separates teams is how long the failures hurt.

Break recovery time into its parts and each one points at something to build:

  • Time to detect: alerts on the symptoms users feel, so the page fires before the support tickets arrive.
  • Time to declare: the cheap declaration from the last section.
  • Time to mitigate: a way to stop the bleeding before you understand the cause. Roll back the deploy, flip the feature flag, shift traffic away from the bad region.
  • Time to resolve: runbooks, dashboards, and enough context that the on-call engineer doesn’t start from zero.

Mitigation deserves the most investment, because it separates recovery from diagnosis. If rollback is one command and two minutes, most incidents become two-minute incidents, and the root cause can wait for daylight. Fast rollback pays off well beyond incidents, too. It’s one of the guardrails that make it safe to let trivial changes ship without a reviewer.

What fast-recovery teams build

When recovery time is the goal, teams start building different things.

They automate the responses they keep doing by hand. If the fix for a backed-up queue is always “add consumers,” that becomes an autoscaling rule, and the incident shrinks to a blip on a graph.

They streamline communication. An incident commander role means the person debugging isn’t also writing status updates. A template for the status page means nobody is drafting customer-facing prose from scratch at 3 a.m.

They end every incident review with the same question: what would have made this faster? The answer becomes a ticket, and the ticket gets prioritized like any other work. A year of those tickets adds up to a system where most incidents are short and boring.

Fewer incidents is an outcome

None of this means incidents don’t matter. You should still ship as fast as possible, but not faster, and a team that keeps causing the same outage has a quality problem that recovery speed won’t hide. Fewer incidents is a fine result. It’s a bad target.

Compare two teams. One declares two incidents a quarter, and each runs six hours. The other declares twenty, and each is over in fifteen minutes. The first team has the better scorecard. The second team’s customers had the better quarter.