Build1 publisher3 min readPublished
Apollon Labs' Greek benchmark finds gpt-oss-20b inventing 48 words per 1,000
Apollon Labs' Greek benchmark found gpt-oss-20b making up 48 words per 1,000, against zero for six of the 15 models tested. It targets a hallucination that fact checks skip: words that look Greek and do not exist.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The benchmark scores answers to 100 short Greek prompts against two fixed lexicons, FrequencyWords' 132,681 subtitle words and the Hunspell el_GR dictionary.
- Every word missing from both lexicons went to a native Greek speaker at Apollon Labs, and the ranking counts only the words that reviewer judged invented.
- Outside gpt-oss-20b, most invented words were real Greek words with a single broken accent, inflection or spelling.
- Claude Opus 5 produced 16 words outside the lexicons, only one of them invented, with real terms such as the Greek for cytosine and perlite making up the rest.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Above about 99.5% the deterministic lexicality score cannot rank models, so on its own it can only screen out weak models.
- cost Porting the method to another language or to technical text means paying a native speaker to rule on each word the lexicons miss.
- exposure A pipeline that screens generated Greek with an LLM will pass some broken forms and reject some real words, as Claude did while helping build this benchmark.
- decision Teams running open weights locally for Greek output now have a per-1,000 rate to choose on, and gpt-oss-20b's is 16 times the worst Claude figure.
Each answer gets three scores, and no model computes any of them. Lexicality is the share of its Greek words found in either lexicon [3]. Greekness is the share of its letters that are Greek, a check on whether the model answered in Greek without being told to [4]. Meaning is whether the answer contains at least one expected keyword. That check exists to catch fluent nonsense [5]. The same answer gets the same score on any machine [6]. It is careful work. Apollon Labs dropped the LLM judge because, the team wrote, "a judge that speaks Greek no better than the models under test can't grade them" [7].
Near the top of the table, the deterministic score runs out of resolution. A lexicon cannot tell an invented word from a rare real one [8]. The authors say a lexicality score above about 99.5% is mostly lexicon noise [14]. For Claude Opus 5, the lexicons flagged 15 real words for every invented one [1]. Ranked on raw lexicality, a model that uses rare, precise vocabulary lands below one that plays safe [14]. So the top of the ranking rests on a native speaker's verdicts, and the post does not report a second rater or an agreement figure for them [8].
Apollon Labs tried handing part of that triage to Claude, its coding partner on the build, and it got things wrong in both directions [16]. Early on it flagged three real words as suspicious, among them the modern neologism προτεραιοποίηση (prioritisation) [16]. Later it accepted the broken form θυμόντουσε as real and marked seven more broken forms as uncertain [16]. "An LLM is not a safe judge of Greek, and that includes the one that helped build this," the team wrote [18].
Most of the errors being counted are morphology errors. DeepSeek R1 wrote the correct imperative Ξεβγάλτε where Haiku 4.5 wrote Ξεβγάλε, so the word itself was not the hard part [12]. Gemini 3 Flash described waves with φλάφισμα, a word that does not exist. The real word is θρόισμα (rustle) [1]. A fact check has nothing to test in either sentence. "Hallucination benchmarks usually check facts. This one checks the words themselves," Apollon Labs wrote [17]. The outlier is gpt-oss-20b. Its 48 invented words per 1,000 are 16 times the highest Claude rate of 3, and coinages such as τρικυδές are not near misses [10][2].
The greekness score sent the team down one wrong path. Gemini models scored 97.7 to 97.8 against 99.9 for GPT, and the first guess was English leaking into the answers [15]. Gemini 3.1 Pro used 12 Latin-script words in 100 answers, the same as GPT-5.5 [15]. The gap came from two empty answers per Gemini model, each with zero Greek letters [15]. If greekness is averaged per answer, two empty answers out of 100 cost 2 points, against a measured gap of 2.1 to 2.2 [3].
All 15 models ran task version v7, with thinking off where the API allowed it and a 1,000-token cap [9]. For the tiers to carry over to production, the traffic has to look like 100 short everyday prompts [2]. I'd expect long technical answers to push more rare real words outside the lexicons. Each of those needs a human verdict before it counts for or against a model [8].
What to watch
- Whether Apollon Labs publishes a second native-speaker rating, or an agreement figure, for its out-of-lexicon verdicts.
- Results with thinking enabled and no 1,000-token cap, where longer answers would put more rare real words in front of the lexicons.
- Whether gpt-oss-20b's 48 per 1,000 holds on longer technical Greek prompts outside the 100-prompt set.