← Back to articles

SOFTWARE TESTING · AUTOMATION

14 min read

The flaky test is telling you something. Usually you don’t want to hear it.

A test that fails one run in twenty is not a nuisance to be retried. It is the only evidence you have of a real defect, and retrying it is how teams delete that evidence on purpose.

Every automated suite of a certain age has them: the two or three tests that pass most of the time and fail occasionally, for no reason anyone has established. The team learns their names. Someone adds a retry. The build goes green and everybody moves on.

The uncomfortable part: a test that fails intermittently is usually reporting a real intermittent defect. Automatic retry does not fix it. It suppresses the report and ships the defect.

1. Where flakiness actually comes from

In my experience the cause is almost never mysterious. It falls into one of five categories, and four of them are the test’s fault.

  • Waiting on time instead of on a condition. sleep(3) is a bet that three seconds is always enough. On a loaded CI runner it is not. This is the single largest cause.
  • Shared mutable state. Two tests read and write the same record, so the result depends on execution order — and on whether they overlap when run in parallel.
  • Test data that is not isolated. A test that assumes a clean database passes alone and fails in a suite.
  • Assertions on unordered data. A query without ORDER BY returns rows in whatever order the engine chose today.
  • A genuine race in the product. The fifth category. Rarer, far more serious, and the one the other four are hiding.

2. Why retry is the wrong response

Consider a checkout that occasionally double-charges under concurrent requests. The defect appears in maybe one run in fifteen. Your suite catches it — that is the system working. Configure two retries and the failure disappears from the report entirely.

# What retry actually computes
p_fail_once      = 0.07          # the defect shows in 7% of runs
p_fail_3_times   = 0.07 ** 3     # = 0.0003

# The defect rate did not change. The reporting rate fell 200x.

Nothing about the product improved. You reduced the probability of being told.

3. What to do instead

Treat a flaky test as a defect report with an incomplete reproduction, because that is what it is.

  • Measure it. Run it a hundred times in isolation and record the failure rate. A test that fails 1 in 100 alone and 1 in 5 in the suite is telling you about interference, not about itself.
  • Quarantine, do not retry. Move it out of the gating suite so it stops blocking, but keep it running and visible. Retry hides it; quarantine keeps the debt on the books.
  • Fix the waiting first. Replace every fixed sleep with a wait for a specific condition. In most suites this alone removes the majority of flakiness.
  • Make each test create and destroy its own data. Independence is what makes parallel execution possible at all.
  • If it survives all of that, suspect the product. By elimination you now have a strong signal, and a concurrency defect found in CI is enormously cheaper than one found by a customer.
A rule I hold to: if a failure is going to be dismissed without investigation, the test should be deleted rather than retried. A suite nobody believes is worse than a smaller suite that is trusted, because it consumes the same runtime and produces no decisions.

4. The organisational failure underneath

Flakiness is rarely a technical problem in isolation. It persists because investigating it is slow and unrewarded, while adding a retry takes one line and turns the build green before the stand-up. The incentive points the wrong way.

The counter-measure is to make flakiness visible as a number. Publish the failure rate per test over the last thirty runs. Once a team can see that four tests account for most of the noise, fixing them becomes a small, finite piece of work rather than an unbounded chore. That reporting is worth building before any individual fix.

5. The claim you are actually defending

The purpose of a suite is not a green tick. It is the ability to say: we ran these checks, against this build, under these conditions, and here is what we found. Every retry weakens that sentence, quietly, and by the time it matters nobody remembers how it got weak.

The Flaky Test Is Telling You Something — Ibrahim Kenia