Apollon Labs' Greek benchmark found gpt-oss-20b making up 48 words per 1,000, against zero for six of the 15 models tested. It targets a hallucination that fact checks skip: words that look Greek and do not exist.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence50
USENIX Security 2025 researchers found 19.7% of packages suggested by 16 LLMs were fake, and 43% of those names recurred on every re-run. Names that repeat can be registered ahead of time, so a team has to vet a suggested dependency before installing it, even when the install succeeds.
Reality
- Evidence58
- Adoption20
- Hype gap+25
- Incentives
- Insufficient
- Confidence50
Engineers on a freight-forwarder sales pipeline swapped an LLM web-search step for a Google Places lookup and lifted website completeness from 34% to 81%. The author credits the gain to the normalisation and match-scoring layer built around the API.
Reality
- Evidence35
- Adoption10
- Hype gap+10
- Incentives
- Insufficient
- Confidence40
Earn an Honest Dollar's benchmark found one prompt line telling models to return null cut invented fields from 70.7% to 20.2% of missing-field answers. The result comes from one synthetic run, and with a fifth of answers still guessing, scraper output still needs a downstream check.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives65
- Confidence45
Researchers found an auto-displayed AI answer cut 'I don't know' responses from 35 percent to 1 percent, though the model was mostly wrong. A 10-cent penalty for wrong answers still left abstention at 7 percent, so review tools need more than an abstain button.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence50
A developer's free resume builder bans its AI from adding nine named kinds of fact and rejects any rewrite that changes the bullet count. The live check covers only the count, so a claim invented in words inside a single bullet still reaches the user, who is told to read every result.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives35
- Confidence45
CNN reports that a Special Operations Command analyst used a chatbot to fuse open-source and classified signals intelligence, then used it again to turn the wrong answer into the summary that moved up the chain.
Perspective Coverage
7 publishers
- Builder
- Builder 33%
- Operator
- Operator 51%
- Investor
- Investor 16%
Reality
- Evidence42
- Adoption48
- Hype gap+25
- Incentives60
- Confidence45
A study of 14 chat-tuned models found their maximum softmax probabilities overconfident everywhere and uncorrelated with task accuracy, while those same scores still sorted correct answers above wrong ones well enough to drive selective abstention.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+8
- Incentives32
- Confidence58
Enhanciar builds a wiki page per service at ingest, makes the model cite those pages, then checks each claim against the cited file before the answer ships. Claims it cannot judge are tagged unverifiable.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+15
- Incentives80
- Confidence40
An autonomous coding setup on a Mac mini splits reading from writing into two agents with separate tool lists. Its operator reports guessed-API failures falling from about one task in five to one in forty.
Reality
- Evidence30
- Adoption10
- Hype gap+25
- Incentives
- Insufficient
- Confidence45
Five frontier models ran a 295-item failure corpus bare and then wrapped. Across the four that accepted the wrapper, confident errors fell from 25.8% to 7.4% and correctness fell further, from 43.9% to 21.3%.
Reality
- Evidence45
- Adoption12
- Hype gap+10
- Incentives60
- Confidence38
A writer of reporting-tool cookbooks put hundreds of developer questions to current chat models and then tested the answers in a lab. The failures he describes cluster into five patterns. Validators clear most of them.
Reality
- Evidence38
- Adoption15
- Hype gap+14
- Incentives72
- Confidence55
In a dev.to account of self-reinforcing memory loops, an agent's own hedged inference is captured as a flat fact, retrieved a week later as context, and generalised until the store recommends replacing the cache layer.
Reality
- Evidence33
- Adoption
- Insufficient
- Hype gap+20
- Incentives40
- Confidence50
The builder shipped 30 Roblox guide sites on Next.js and Vercel in a month, then audited them and found invented code tables across most of the network, plus a canonical tag pointing at a domain he never registered.
Reality
- Evidence42
- Adoption22
- Hype gap+25
- Incentives40
- Confidence45
Rulestack says only one of the five was the kind of error a string matcher could catch, and the first matcher it wrote cleared the draft anyway, because 76 turns up 1,343 times in the ledgers it searched.
Reality
- Evidence52
- Adoption15
- Hype gap−8
- Incentives58
- Confidence45
A one-dimensional probe finds structural impossibility in the hidden state of instruction-tuned models from 1.7B to 70B parameters. The trained refusal pathway that guardrail work tunes reads an axis about 85 degrees away from it.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−12
- Incentives40
- Confidence52
Gizmodo asked five AI systems whether a viral image of the 9/11 rescuer Welles Crowther was real. Each described the fabrication as a photograph of some kind, and what caught it was a reverse image search.
Reality
- Evidence62
- Adoption38
- Hype gap+20
- Incentives55
- Confidence58
When registry-mcp measured the counterfactual its README had been asserting, claude-sonnet-5 mostly declined to answer rather than inventing filings. That moves the argument for the tool onto provenance and coverage.
Reality
- Evidence56
- Adoption
- Insufficient
- Hype gap−18
- Incentives68
- Confidence61
Slopsquatting stopped being a thought experiment. The only thing standing between an AI suggestion and an install was a reviewer who happened to check the registry page.
Reality
- Evidence26
- Adoption12
- Hype gap+38
- Incentives86
- Confidence33
Engineers at AVIC Chengdu warn that language models invent radar ranges, payload figures and fatigue limits with full confidence, and that fabricated threat data lands upstream of design decisions.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+24
- Incentives58
- Confidence36