Skip to content

Build1 publisher3 min readPublished

One wrong punctuation mark cost Claude models more accuracy than misspelling 70% of the words

Misspelling 70% of a prompt's words left Claude models' scores unchanged across about 4,900 test sessions, but one wrong punctuation mark cost 8 to 23 points. Both breaks that stuck erased the line between instruction and data, so delimiters deserve the review time that spelling gets.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying One wrong punctuation mark cost Claude models more accuracy than misspelling 70% of the words
Generated illustration

What happened

  • Removing the closing quote from a word-counting task made every model count three extra instances of "report" and answer 8 instead of 5.
  • With the colon dropped before a word list, Sonnet and every local model treated "reply" as data and returned its plural alongside the others.
  • A comma-less ticket query with two possible readings resolved correctly on all five Claude models because the sentence opened with "For Lee's standup".
  • On dictated input, Sonnet read "two fifty" as $250 in six of six trials and answered 375,000 instead of 3,750.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure An evaluation harness that extracts only the final answer will accept the wrong count even when the model's own reply names the missing quote.
  • constraint The two cheap remedies a team would try first left the structural failures in place, so the repair has to live in the code that assembles the prompt.
  • decision Prompt review time is better spent on the quotes and colons that bound inserted data than on spelling, because context repaired the other breaks for the Claude models.
  • cost Teams putting dictation in front of small local models need a normalisation step for spoken numbers and list markers before the prompt is sent.

The two breaks that context did not repair both sat on a delimiter [9]. A closing quote ends the scope of "inside the quoted paragraph." A colon before a list ends the instruction and starts the data. Remove either one and the prompt still reads as English, with the boundary in a different place [10][12]. The other breaks, such as a comma moved around "except", were recovered from context by every Claude model [4][9].

Detection did not produce correction. Opus 5.5 wrote "the closing quote is missing, so I counted all" and still returned the wrong number [11]. It diagnosed the bug correctly and shipped it anyway. Raising effort from low to medium made no difference for any model on any task [18]. "Structural breaks are not a reasoning problem, so more thinking does not fix them," the author wrote [19]. A prepended line telling the model to expect spelling, grammar and punctuation mistakes had zero effect on Opus 5, Haiku and Sonnet, and added a few points on Opus 5.5 and Fable [20].

The spelling result holds under one condition. The script misspelled 35% or 70% of the words in the long prompts but never touched the words the answer depended on [3]. The finding covers noise around the key terms. It transfers to a production prompt only if those terms arrive intact. Scoring was lenient: a reply counted as correct if the ideal answer appeared anywhere in it, explanation included [6]. A harness that parses one bare field is stricter. The hard set has nine prompts [3]. Weighted equally, each is about 11 points of a column, so a loss of 8 to 23 points [8] is roughly one to two prompts failing [1]. That matches the two breaks context did not repair [9].

The seven local models, run through Ollama, took the damage from dictated input [5]. On the voice-to-text profile, gemma4 fell 34 points and granite 37 [2]. The author traced it to spelled-out numbers and, again, the missing colon before a list; every local model scored 0% on the dictated list task [15]. The Claude models mostly absorbed dictation. Beyond Sonnet's run of misreads, Fable made the same "two fifty" error once, and Haiku turned "extension two zero one" into 2001 once [16]. "These are not model bugs. Spoken numbers are ambiguous, and the model has to guess," the author wrote [17].

The test introduced each break deliberately and did not include prompts built from templates [1][4]. I think templated prompts are where the delimiter finding applies most. A template is the code that puts the quote or colon between the instruction and the inserted text, and a bug there repeats on every call. I'd check what a template emits around inserted data before spending any effort on the spelling of what users type. The prompt that began "shop sellin pens 3dolar each" scored 100% on the Claude models [21].

What to watch

  • A rerun in which the misspelling script is allowed to hit the words the answer depends on, which this test deliberately spared.
  • Scores under strict single-field answer parsing in place of the test's check that the ideal answer appears anywhere in the reply.
  • The same delimiter breaks produced by a real templating layer, such as inserted text that carries its own quote marks.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories