Meta, Duke and UC Davis researchers lifted Gemini 3 Flash from 46% to 62% on Olympiad math by searching for a better harness around the same model. That puts the wrapper on an operator's budget ahead of a bigger model, at least for math.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
Ridge Security's 96-test pen-test benchmark had Claude Opus 4.6 reach 63% coverage at $217 a run, against 52% for Gemini 3 Flash at about $5.42. Ridge argues tooling matters more than the model, yet its published runs held the tooling fixed and varied the model.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+40
- Incentives70
- Confidence40
Apollon Labs' Greek benchmark found gpt-oss-20b making up 48 words per 1,000, against zero for six of the 15 models tested. It targets a hallucination that fact checks skip: words that look Greek and do not exist.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence50
A 73-page preprint evolved instructions that jumped between coding agents and wrote themselves into the file that becomes the next system prompt. A short warning nearly stopped transmission.
Perspective Coverage
4 publishers
- Builder
- Builder 52%
- Operator
- Operator 39%
- Investor
- Investor 9%
Reality
- Evidence68
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence65
Antigravity's free tier includes Gemini 3.1 Pro and Claude Opus 4.6 with unlimited tab completions. XDA reports Opus burns credits about four times faster than Gemini, and Google moved the free tier onto credits in March 2026.
Reality
- Evidence36
- Adoption
- Insufficient
- Hype gap+12
- Incentives58
- Confidence40
TypeSafe's Jev answers typed questions in one forward pass with no token stream, and a dev.to benchmark shows that most of its 14x decision-latency lead over two chat models came from how those models were called.
Reality
- Evidence58
- Adoption10
- Hype gap+12
- Incentives55
- Confidence45
Vercel's September Production Index shows the gateway's average price per token down 23.2% in August, a third straight decline, and the median heavy-usage team paying 7.6% less, so most of the saving came from switching models.
Reality
- Evidence58
- Adoption80
- Hype gap+14
- Incentives70
- Confidence64
The company's own blog post is the only account of the run, and an Adelaide University researcher who credits its long-horizon design still warns that open-ended worlds make causes harder to isolate and runs harder to compare.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+40
- Incentives74
- Confidence44
The Kotlin Benchmark grades agents on 105 verified repository tasks. Its token column shows setups that solve within a few tasks of each other burning between 66,000 and 777,000 tokens per fix.
Reality
- Evidence62
- Adoption30
- Hype gap+22
- Incentives60
- Confidence58
Parse 5 returns Markdown and bounding boxes from a model small enough to serve on Azure or SageMaker, and it lands eleven points below the ParseBench leader. The 8K context window is the number to check first.
Reality
- Evidence45
- Adoption15
- Hype gap+30
- Incentives72
- Confidence42
Databricks argues agents fail on charts at retrieval rather than at reasoning, and it tested the idea with two indexes over the same 16,000-page corpus differing only in how figures were represented.
Reality
- Evidence56
- Adoption
- Insufficient
- Hype gap+32
- Incentives88
- Confidence46
Premium requests are now labelled legacy and AI Credits bill at one cent each. Annual Pro and Pro+ seats keep the old meter, so one org ends up budgeting in two units with a four-to-one conversion.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+22
- Incentives55
- Confidence36
MLX LoRA has no per-example weight field, so one builder encoded his curriculum as duplicate lines. A dedup key on the last 200 characters deleted 38,988 of them before training.
Reality
- Evidence63
- Adoption14
- Hype gap−9
- Incentives32
- Confidence57
A preprint from ServiceNow AI Research and university co-authors reports that a plain web agent with a longer horizon matches or beats AWM, ASI and ReasoningBank, often on fewer tokens.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives45
- Confidence42
An Anthropic and EPFL preprint shows plain-language goals hopping agent to agent through persistent files, and a one-paragraph warning in the system prompt stopping nearly all of it.
Publishers:startupfortune.com
Reality
- Evidence55
- Adoption22
- Hype gap+15
- Incentives58
- Confidence42
Self-propagating payloads did move between agents through editable soul files, but one inoculation paragraph held against 150-plus optimized strains, and nothing propagated in the wild.
Reality
- Evidence66
- Adoption14
- Hype gap+18
- Incentives60
- Confidence55
Google Research reports Gemini-3-Pro and GPT-5 encode 95-98% of tested facts yet fail to recall 26-34% of them, moving the fix from pretraining scale toward post-training and inference.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives68
- Confidence48