Build1 distinct publisher3 min readUpdated
An agent-built resume generator passed generation, rendering and ATS checks while three product requirements were still wrong. None of the three tripped a script.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
A developer building a private "facts to prose" resume generator with a Cursor agent reported that generation, rendering and ATS checks all stayed green while several product requirements were still wrong [1][2][3]. That is the failure mode worth naming for anyone shipping agent-written tooling: the gates confirmed that artifacts were produced, not that the specification described the artifact anyone wanted [4][6].
The setup is ordinary. Structured career claims go in, recruiter-facing PDFs come out [1]. Implementation QA meant the pipeline ran end to end and the automated checks passed: generation succeeded, PDFs rendered, and ATS scripts asserted page counts, required sections and a few structural rules about separators and headings [4]. Those checks were not theatre. According to the author they caught broken builds and regressions he did not want to ship [5]. They simply had no view on whether the spec matched the output he wanted [6].
Three misses illustrate the gap, and none of them failed the scripts at first [7]. The workflow produced one-page and two-page variants for the same application; both passed check:ats and both stayed inside their page limits [8]. The longer version reused the same condensed evidence selected for a tight one-page fit and then filled the remaining space with additional facts, so the concise bullets never expanded into fuller evidence [9]. Page count was correct; the shape was wrong [9].
Second, a Program Lead role was valid in the data with correct dates, employer and title, and rendered with a Full-time employment label, which read like a sequential primary job when the work was actually concurrent [11]. Renaming it to "Concurrent program" was a product fix, and no ATS script flagged the old wording [12].
Third, the contact block and Skills sidebar shared a column edge, but the Skills heading sat a few points lower than Experience, so the two-column header row looked crooked while every section and separator rule still passed [13]. Add it up: three documented product defects, zero flagged by the gates [1].
The most instructive detail is the one about fonts. Mismatched typography could have faked the same green result by enlarging text on the longer PDF; unifying the shared typographic scale and margins removed that shortcut, so a two-page count had to come from the evidence itself [10]. That is the general shape of a weak gate. If a metric can be satisfied by a change nobody asked for, the metric is measuring production, not intent. The fix was not a better assertion but a constraint that made the cheap path unavailable.
The author's diagnosis of why this is newly acute is worth taking seriously. When he wrote this tooling by hand, implementation and requirements review were hard to separate, because choosing a font size or cutting a bullet forced a product judgment in the same session as the code change [15]. Manual work was doing requirements QA by accident; an agent absorbs that ambiguity instead, and incomplete specification arrives dressed as a finished product [15][16].
The remedy he offers is procedural, not tooling: force the vague request into concrete behavior before implementation, naming what changes, what stays invariant and what counts as done [17]. Note that the workflow had encoded structural correctness but never named a human visual-acceptance criterion, the plain question of whether he would send it [14]. Until that question is written down somewhere a reviewer must answer, a green build is evidence about the machine and not about the product.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The author used a Cursor agent to implement a private facts-to-prose resume generator: structured career claims in, recruiter-facing PDFs out.
Generation, rendering and ATS checks all stayed green, and the PDFs looked plausible.
Several product requirements were still wrong despite the passing checks.
Implementation QA in this workflow meant the pipeline ran end to end and automated checks passed: generation succeeded, PDFs rendered, and ATS scripts asserted page counts, required sections and a few structural rules about separators and headings.
The checks were real and caught broken builds and regressions the author did not want to ship.
The checks did not answer whether the specification described the resume the author actually wanted.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but unverifiable single-author account
The account is specific and internally coherent: it names the check suite's assertions, three concrete defects, and the fixes applied. But everything rests on one self-reported post about a private project, with no repository, diffs, artifacts, screenshots or second observer, and no measurement of frequency or generality.
One private single-practitioner workflow
The only adoption signal is the author's own disclosed use of a Cursor agent on a private, personal-scale resume pipeline, plus the workflow ladder they say they follow. No teams, organizations, downloads, deployments or third-party usage are reported.
Deflationary framing, modestly under-claimed
The piece argues against over-reading a green pipeline and explicitly credits the checks with catching real regressions, so it does not inflate its own findings. If anything the framing is narrower than the material: a reusable failure mode about agent output passing structural gates is presented as one person's resume-tooling anecdote, with no attempt to claim generality it cannot prove. Slightly negative rather than zero.
Practitioner self-publishing, no disclosed vendor stake
This is a personal dev.to post that showcases the author's engineering judgment and links back to an earlier piece of their own, so there is a visible personal-brand incentive. There is no disclosed sponsorship, no product being sold, and the one named vendor (Cursor) is mentioned only as the tool used, without promotional framing or a commercial relationship stated in the source.
Plausible testimony, low corroboration
Confidence is limited by structure rather than plausibility: one publisher, one author, one private project, no artifacts, and the two most transferable claims are explicitly generalizations from a single anecdote. The descriptive detail is concrete enough to trust as testimony about what happened in that workflow, but not enough to support inference about how common the pattern is.
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
build
48 startups, 4 known by name, 28 recommended by category1 distinct publisher
build
Thirteen tasks green, then "give up (Recommended)" on the one that needed understanding1 distinct publisher
build
An agent built and deleted a prod stack. The alert fired on time and changed nothing1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 17, 2026