Skip to content

Build1 publisher3 min readPublished

Telling extractors to return null cut invented fields from 70.7% to 20.2% in a synthetic benchmark

Earn an Honest Dollar's benchmark found one prompt line telling models to return null cut invented fields from 70.7% to 20.2% of missing-field answers. The result comes from one synthetic run, and with a fifth of answers still guessing, scraper output still needs a downstream check.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Telling extractors to return null cut invented fields from 70.7% to 20.2% in a synthetic benchmark
Generated illustration

What happened

  • The test used 42 pairs of synthetic pages across seven page types, and each page missing the answer kept a decoy, such as an old price marked "Was $493.00".
  • GPT-6 Luna, asked to verify returned values against the page, caught 38 of 49 made-up values and wrongly rejected none of 47 correct ones.
  • Checking 126 unique page-and-value pairs with GPT-6 Luna cost $0.0049 in the reported run.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Teams that write their own extraction prompts can add the null line at the cost of one sentence; on this run it removed about 71% of invented missing-field answers.
  • constraint Teams buying extraction from a hosted API get less from the line: Firecrawl's free tier still invented 66.7% of missing fields with it, near the bare models' uninstructed 70.7%.
  • cost A GPT-6 Luna verification pass costs about 3.9 cents per thousand values at the reported price, yet it let 11 of 49 invented values through, so buyers still have to decide how much error they accept.
  • exposure Fields with close neighbours in meaning, such as preparation time beside resting time, got past the checker, so those fields need their own validation rules or human review.

The line is short enough to paste into any extraction prompt: "Use null for any field whose value is not on the page. Do not guess" [1]. The published percentages match the raw counts. Without the line, 405 of 573 answers were invented, or 70.7%. With it, 116 of 574 were, or 20.2%, a relative drop of about 71% [5][1][2]. These are response counts pooled across 16 models. Each contestant ran once, on pages built for the test [5][6]. Earn an Honest Dollar published the results on September 27th [2]. Only pages where the requested field was absent were scored [4].

Each of the 42 page pairs had one version with the answer and one without it, with a decoy left in place [3]. The decoys included an old price labelled "Was $493.00", a fact-checker placed near the author field, and an update date that could pass for a publication date [3]. Without the null line, every model reported $493 as the current price, a bug anyone who has maintained a price scraper has already filed [7]. With the line, one model still did [7].

For the 71% to carry over, production pages need decoys shaped like these. The model has to respond to the instruction the way these 16 did on average. The instruction also has to reach the model intact. I would add the line to any extraction prompt I own, since it costs one sentence. I would not drop a downstream check on the strength of one synthetic run.

The hosted services put the line behind someone else's pipeline. Firecrawl, tested on a free tier with the instruction, made up 24 of 36 missing fields, and according to the benchmark all 24 were copied from a decoy [8][10]. Its rate of 66.7% is close to the 70.7% the bare models posted with no instruction at all [3][1]. ScrapeGraphAI made up 7 of 31, or 22.6%, and ScrapingBee 16 of 36, or 44.4% [9][3]. ScrapingBee has no prompt or schema field, so the null rule went into each field description [10]. None of the services was run without the instruction, so there is no before figure for them. The benchmark's own caveats say paid plans may behave differently [10].

The checker step follows from what the site sells. Earn an Honest Dollar runs a marketplace where providers list services and execution endpoints for other agents to find and buy. It argues those buyers need a way to judge quality without inspecting every answer [15]. In the benchmark, a second model checked whether each returned value was supported by the page. GPT-6 Luna caught 38 of 49 made-up values and rejected none of 47 correct ones [11]. Jev 1.13 caught 23 of 49 and rejected none of 48 correct ones [12]. The catch rates are 77.6% and 46.9% [4]. According to Earn an Honest Dollar, checking 126 unique page-and-value pairs with GPT-6 Luna cost $0.0049, about 3.9 cents per thousand checks [14][5].

Suppose GPT-6 Luna's catch rate held on the extractors' output. Then the 20.2% that still guessed would fall to about 4.5% of missing-field answers, roughly 26 of 574 [6]. That estimate assumes the checker fails independently of the extractor. The misses the benchmark describes point the other way. The checker passed near-meaning errors such as cooking or resting time returned as preparation time [13].

What to watch

  • A rerun of Firecrawl, ScrapeGraphAI and ScrapingBee on paid plans, or with a no-instruction baseline, showing whether the hosted APIs gain from the null line at all.
  • Repeated runs or live-site pages from Earn an Honest Dollar, which would test whether the 71% relative drop survives outside one synthetic run.
  • Whether Earn an Honest Dollar's marketplace exposes benchmark scores or checker results to buyer agents when they pick a service.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories