Skip to content

Build1 publisher3 min readPublished

Green pipeline, wrong product: the checks proved the PDF rendered, not that anyone wanted it

An agent-built resume generator passed generation, rendering and ATS checks while three product requirements were still wrong. None of the three tripped a script.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Green pipeline, wrong product: the checks proved the PDF rendered, not that anyone wanted it
Generated illustration

What happened

  • The author used a Cursor agent to implement a private facts-to-prose resume generator: structured career claims in, recruiter-facing PDFs out.
  • Generation, rendering and ATS checks all stayed green, and the PDFs looked plausible.
  • Several product requirements were still wrong despite the passing checks.
  • Implementation QA in this workflow meant the pipeline ran end to end and automated checks passed: generation succeeded, PDFs rendered, and ATS scripts asserted page counts, required sections and a few structural rules about separators and headings.
  • The checks were real and caught broken builds and regressions the author did not want to ship.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer building a private "facts to prose" resume generator with a Cursor agent reported that generation, rendering and ATS checks all stayed green while several product requirements were still wrong [1][2][3]. That is the failure mode worth naming for anyone shipping agent-written tooling: the gates confirmed that artifacts were produced, not that the specification described the artifact anyone wanted [4][6].

The setup is ordinary. Structured career claims go in, recruiter-facing PDFs come out [1]. Implementation QA meant the pipeline ran end to end and the automated checks passed: generation succeeded, PDFs rendered, and ATS scripts asserted page counts, required sections and a few structural rules about separators and headings [4]. Those checks were not theatre. According to the author they caught broken builds and regressions he did not want to ship [5]. They simply had no view on whether the spec matched the output he wanted [6].

Three misses illustrate the gap, and none of them failed the scripts at first [7]. The workflow produced one-page and two-page variants for the same application; both passed check:ats and both stayed inside their page limits [8]. The longer version reused the same condensed evidence selected for a tight one-page fit and then filled the remaining space with additional facts, so the concise bullets never expanded into fuller evidence [9]. Page count was correct; the shape was wrong [9].

Second, a Program Lead role was valid in the data with correct dates, employer and title, and rendered with a Full-time employment label, which read like a sequential primary job when the work was actually concurrent [11]. Renaming it to "Concurrent program" was a product fix, and no ATS script flagged the old wording [12].

Third, the contact block and Skills sidebar shared a column edge, but the Skills heading sat a few points lower than Experience, so the two-column header row looked crooked while every section and separator rule still passed [13]. Add it up: three documented product defects, zero flagged by the gates [1].

The most instructive detail is the one about fonts. Mismatched typography could have faked the same green result by enlarging text on the longer PDF; unifying the shared typographic scale and margins removed that shortcut, so a two-page count had to come from the evidence itself [10]. That is the general shape of a weak gate. If a metric can be satisfied by a change nobody asked for, the metric is measuring production, not intent. The fix was not a better assertion but a constraint that made the cheap path unavailable.

The author's diagnosis of why this is newly acute is worth taking seriously. When he wrote this tooling by hand, implementation and requirements review were hard to separate, because choosing a font size or cutting a bullet forced a product judgment in the same session as the code change [15]. Manual work was doing requirements QA by accident; an agent absorbs that ambiguity instead, and incomplete specification arrives dressed as a finished product [15][16].

The remedy he offers is procedural, not tooling: force the vague request into concrete behavior before implementation, naming what changes, what stays invariant and what counts as done [17]. Note that the workflow had encoded structural correctness but never named a human visual-acceptance criterion, the plain question of whether he would send it [14]. Until that question is written down somewhere a reviewer must answer, a green build is evidence about the machine and not about the product.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories