Build1 publisher3 min readPublished Updated
Coverage Is A Line Counter, So A Coverage Gate Buys You Line-Counting Tests
One engineer's billing flow broke on launch day at 87% coverage. The metric was doing exactly what it was built to do, which is why the fix starts with incentives rather than tooling.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The author writes: 'I had 87% coverage, and we still broke the billing flow on launch day. Not because of a gap in the percentage. Because 87% was covering the wrong things. The tests were written to pass a gate, not to catch a failure.'
- At 87% line coverage, 13% of lines were unexecuted; the author attributes the launch-day billing failure not to that 13% but to what the covered lines were checking.
- Coverage percentage tracks which lines of code were executed during a test run. If a line ran, it counts as covered. The source states that this is the complete definition.
- Coverage does not measure whether the test asserted anything meaningful about the line, whether both branches of a conditional were exercised, or whether the specific inputs that cause failures were ever tried.
- A test that calls a payment function and checks 'assert response is not None' covers the same lines as a test validating the transaction ID, amount, currency, error code and retry behaviour; the coverage tool treats them identically.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
An engineer writing on dev.to describes shipping a billing flow that broke on launch day while the suite sat at 87% coverage, and argues the failure was not in the unexecuted remainder but in what the covered lines were actually checking [1][1]. That is worth attention because coverage percentage is the answer most teams give when asked whether their testing is working, and the metric records one thing: which lines ran during the test run [3].
That is the whole definition. If a line executed, it counts [3]. Coverage does not know whether the test asserted anything meaningful about that line, whether both branches of a conditional were exercised, or whether the inputs that actually cause failures were ever tried [4]. A test that calls a payment function and checks that the response is not null covers the same lines as one validating transaction ID, amount, currency, error code and retry behaviour. The tool scores them identically [5].
The empirical record, as the post summarises it, is not kind to the metric. Kochhar et al. in 2017 looked at coverage against post-release bug counts across 100 large open-source Java projects and found the coverage of existing test suites had an insignificant correlation with bugs found after release [6]. Inozemtseva and Holmes separately found that line coverage correlates least with real defect detection among the available options [7].
So the gate is a proxy, and Goodhart's Law applies: when a measure becomes a target, it stops being a good measure [8]. The behaviour that follows is not carelessness, it is arithmetic. Under delivery pressure, the cheapest route to a number is the happy path: those tests are quick to construct, run fast, and each one sweeps a high line count because it walks the main execution path [9]. Error handling, retry logic, fallback and timeout behaviour sit in branches that need deliberate setup, and those are precisely where the expensive bugs live [10]. Boundary inputs, nulls, empty strings, Unicode, maximum lengths, each cost a separate test while covering nearly the same lines as one ordinary input [11]. A conditional reads as covered the moment either side executes, so a function that treats admin and guest differently can show green having only ever seen one role [12].
The result the author describes is a suite at 85% where the safe parts are tested three times over and the failure-prone parts carry one shallow test each [13]. The post also claims that teams with high line coverage and low integration validation have been observed producing more than double the production incidents of teams with lower but more balanced coverage; that figure is the weakest link in the argument, since the piece flags it as an assumption rather than a sourced result [14][15].
Strip the metric criticism away and the operational point stands: this is an incentive that management installed, and engineers responding to it rationally [16]. Changing the tool without changing what gets rewarded moves the gaming somewhere else.
What to watch on your own team: whether the error paths, retries and timeouts have tests at all, since coverage will not tell you [10]; whether conditionals have been exercised on both sides rather than counted once [12]; and whether release sign-off still hinges on a percentage. The author's framing is that this is a management decision before it is a tooling decision [15], and that ordering is testable. If the gate stays, the tests will keep being written for the gate.