Skip to content

Build1 publisher3 min readPublished

An AI test suite hit 94% coverage and missed the one branch that mattered

A hand-written mutation survived all 20 generated pytest tests on a small CSV parser. Coverage counted lines executed; nothing in the suite checked the behaviour.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying An AI test suite hit 94% coverage and missed the one branch that mattered
Generated illustration

What happened

  • The author has a small Python library that parses CSV files and extracts course schedules; it reads a CSV, skips header rows, and returns a list of dictionaries, with a function for handling missing values and another for date parsing. It had no tests.
  • The author used MonkeyCode's free model access and free server, and the article discloses it was prepared as part of MonkeyCode's product outreach.
  • The prompt asked the model to write pytest tests for the module covering normal cases, missing values, and invalid dates, and to return only the test code.
  • The model produced 20 tests.
  • Run with pytest and coverage.py, the report said 94% coverage, and line 12, the row.get('Course') check, was the only miss.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer with an untested Python CSV parser asked a free model for a pytest suite, got 20 tests back, and measured 94% line coverage [1][4][5]. He then changed one boolean condition in the source by hand, and every test still passed [8].

The library is about as plain as shipping code gets: open a CSV with DictReader, keep rows where both `Course` and `Time` are present, parse the time string, return a list of dicts [1][6]. `parse_time` returns a time object, or `None` when `strptime` raises `ValueError` [7]. The prompt asked for normal cases, missing values and invalid dates, and for test code only [3]. Under pytest and coverage.py the report came back at 94%, with line 12, the `row.get('Course')` check, as the only miss [5].

The mutation was one clause: `row.get('Course')` became `row.get('Course') or True`, which makes half the guard vacuous [8]. The suite passed and coverage still read 94% [8]. The reason is not subtle. None of the generated tests fed in a row without a `Course` key; they exercised a missing `Time` but never a missing `Course`, and every fixture carried a `Course` column [9]. With the mutant in place, any row missing a course name flows straight into the output [10]. A second mutation did get caught: making `parse_time` return `datetime.now().time()` instead of `None` on invalid input failed `test_parse_time_invalid`, which asserted `None` [11].

So one of two hand-picked mutants died, a 50% kill rate [1] against a 94% coverage number [5], a spread of 44 points on metrics that are not measuring the same thing [3]. Two mutants chosen by the author, applied manually rather than by a mutation runner, is a spot check and not a mutation score [8]. It is still the failure mode you should expect, because the model wrote both the assertions and the fixture files they read from, `sample.csv`, `missing.csv` and `empty.csv` [12]. An input shape the model did not imagine is therefore absent from both sides of the test, and coverage cannot flag it: the missing case is not a line, it is an input. Note also where the two signals landed. The single line coverage reported as unexecuted and the line whose mutant survived are the same line [2].

The author's own reading is that coverage measures lines executed, not behaviours verified, that it is a proxy rather than a guarantee, and that mutation testing is the better check [13][15]. He also rates 20 tests in seconds as a real improvement on zero, as a starting point [14]. Worth knowing before you weight any of this: the piece discloses that it was prepared as part of MonkeyCode's product outreach, and MonkeyCode's free model and free server were what generated and ran the suite [2].

Watch what your acceptance gate actually reads. If generated tests land in CI against a coverage threshold, that threshold is measuring the model's ability to execute lines, which is the thing it is best at. The cheaper tell, before you install a mutation runner: find every `.get()` and every default-value branch in the code under test, then check whether any fixture supplies the absent case. In this instance none did [9].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories