Matthew Boston

How Do You Know You've Fixed a Flaky Test?

May 8, 2025

A flaky test passes most of the time. That’s what makes it flaky. So when you push a fix and the build goes green, you’ve learned almost nothing, because the build was probably going to go green anyway. You know you’ve fixed a flaky test when you can make it fail on demand, name the cause, and run the same reproduction afterward without a single failure.

A green build proves nothing

Say a test fails 2% of the time. Run it ten times with no fix at all and there’s an 82% chance all ten pass. Run it a hundred times and there’s still a 13% chance you see nothing. Most flaky test “fixes” get verified with one CI run, which is the same as not verifying them.

Statistics has a handy shortcut for this, the rule of three. If you run a test n times and see zero failures, you can say with about 95% confidence that the true failure rate is below 3/n. Zero failures in 300 runs puts it under 1%. Zero in 3,000 puts it under 0.1%. One green build puts it under 300%, which is not a useful bound.

Make it fail before you touch it

If you can’t reproduce the failure, you can’t prove you fixed it. The first job is a loop that fails often enough to measure.

Start by running the test by itself, many times:

bash go test -run TestCheckout -count=500 -race ./checkout/

Every test runner has some version of this: a repeat flag, a plugin, or a shell loop. If it never fails in isolation, look between tests. Something another test does is leaking into this one. Run the whole suite in random order and keep the seed. RSpec prints the seed it used and replays it with --seed, and rspec --bisect will narrow an order-dependent failure down to the smallest set of examples that reproduces it.

Then change the conditions. Run with parallel workers. Run on a loaded machine, with another suite going alongside. Freeze the clock at 11:59 PM on the last day of the month. CI runners are slower and busier than your laptop, and a lot of flakes only show up when timing gets worse.

You want to turn a 1-in-500 failure into a 1-in-5 failure. Once it fails that often, you can test a theory in seconds.

Name the cause

A fix you can’t explain is a guess. Most flaky tests land in a handful of buckets:

  • Shared state. One test leaves a row in the database, a value in a global, or a file on disk, and the next test trips over it.
  • Ordering assumptions. The test expects records back in insertion order, and the query has no ORDER BY. The database never promised an order, and one day it stops giving you the one you expected.
  • Time. The test computes “tomorrow” a minute before midnight, assumes every month has 30 days, or reads the real clock twice and gets two different answers.
  • Async work. The test asserts before the background job or the UI render has finished.
  • The network. I’ve already covered why tests that make network requests are flaky by definition, so I won’t repeat it here.

When you find the cause, you should be able to write it in one sentence in the commit message. “The test assumed users came back sorted by id, and the query never asked for that.” That sentence is what makes the fix reviewable, and it’s what tells the next person which bucket to check first.

Retries and sleeps don’t count

A sleep(2) before the assertion changes the odds. The race is still there, and the test will fail again on the day CI is slow enough. Wait on the actual condition instead: poll until the job reports done, with a timeout, or run the async work synchronously in tests.

Automatic retries are worse. A retry plugin with one retry turns a 5% failure into a 0.25% failure, the build goes green, and nobody looks again. Sometimes the flaky test is the only thing telling you about a real race condition in production code. Retry it into silence and you’ve thrown that signal away.

Prove it, then keep watching

Run the same reproduction loop after the fix that you ran before it. If the loop failed 14 times in 500 before and zero times in 500 after, you have evidence. Put both numbers in the pull request.

Then watch it for weeks. Record every test result from every CI run, and count a test as flaky when it passes on retry or flips from red to green on the same commit. A per-test flake rate over time tells you whether the fix held, and it tells you which test to fix next. If the test you fixed shows up on that list again, the first fix was a guess.

I’ve been sounding the alarm about flaky tests for close to a decade, and the re-run button still bothers me most. Every re-run teaches the team that a red build might not mean anything. Measured fixes are how you get back to red meaning broken.