Build1 publisher3 min readPublished
Resume-parsing guide: disclosure team owns the field allowlist, even if another team owns the template
Guidance on dev.to says redacted resumes must be new documents rendered from a team-owned template, with one of three ownership models picked first. That hands the field allowlist to whoever answers for disclosure and leaves the PDF parser as a swappable adapter.
The Engineer · Build desk
What happened
- The guide's pipeline sends private PDF bytes to an extraction worker, lets a model propose a candidate record, and passes only validated fields to the template.
- The model returns a closed schema with nulls for uncertainty and short page-numbered quotes, and it proposes fields without deciding what may be shared.
- Normalization keeps page breaks and collapses repeated spaces inside a line, but leaves names and punctuation untouched before extraction.
- When extraction returns no text, the document goes to OCR or manual review and never reaches the model.
- Logs record identifiers and counts, with resume content kept out of them.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Assigning the field allowlist becomes the first build decision for candidate-summary features, ahead of parser selection, because the parser can be replaced behind its adapter.
- cost Teams that let operations edit layouts take on a leakage test run for every published template version before it goes live.
- exposure Code copied from the guide can hand the model interleaved two-column text that still clears the page-cited evidence check.
- constraint Supporting customer-supplied formats requires an isolated renderer that fails closed, so some customer requests will end with no document at all.
The record the guide extracts still holds the personal fields. Its CandidateRecord type carries displayName, email and phone beside skills, experience and page-cited evidence [18]. Under application ownership, the renderer takes "a narrow type instead of an open bag of resume fields," and a template change passes through code review, fixtures and deployment beside its disclosure policy [5]. A field left out of that type cannot reach an interviewer's copy, and a fixture test can confirm it [5][2].
The guide asks for a conforming PDF implementation behind a small interface that returns page, x, y and text for each span, and keeps parser-specific objects out of the privacy-sensitive service [9]. Swapping parsers then means rewriting one adapter [9]. Changing the template owner changes who can widen what gets disclosed, and the guide puts that choice first [2].
Reading order stays on the parser side, and there the guide's own code has a gap. It warns that sorting by y alone can interleave two columns [10]. Its spansToText function sorts spans by descending y, then x, adds each span to the first line whose opening span sits within a default tolerance of 2 on the y coordinate, and joins each line left to right [11]. Nothing in that grouping looks at x, so it runs across the full page width [11]. The code does not follow its own warning [10][11].
On a two-column resume, a left-column skill and a right-column employer at the same height become one line of text [16]. Validation rejects evidence quotes absent from the normalized text [15]. That text is the interleaved output, so a quote copied from it passes [16].
The other two ownership models buy flexibility with process. Operations editors get named placeholders such as candidateLabel and skills. Each template version is published and leakage-tested before activation, and the guide calls the review burden real [6]. Recipient-owned templates are untrusted input, rendered in an isolated worker with unknown placeholders rejected and remote resources disabled [7]. A customer who asks for email in an anonymous profile gets a failed render [7]. The guide's decision rule covers all three: "the team accountable for disclosure owns the field allowlist, even when another team owns the visual layout" [8].
The evidence here is one tutorial's design argument, with no incident data or test results attached. On limits, it tells readers to cap bytes, pages, extracted characters and processing time before the model, using their own workload measurements [13]. The ownership argument follows from the types it defines: the parser's output is page, coordinates and text, and the renderer's input is whatever the allowlist permits [9][5]. The line grouping in spansToText needs a column-detection step before its evidence check can confirm reading order on a multi-column resume [16].
What to watch
- A revision of the guide's spansToText that detects columns before grouping lines, which would close the gap between its reading-order warning and its code.
- Published leakage-test fixtures for the operations and recipient ownership models. The guide's design argument currently comes with no test results, and fixtures would let readers check a template for leaks themselves.