Cloudflare released two Jev-API-compatible decision models, Clef and Clef-flash, on Workers AI and as Apache 2.0 weights on Hugging Face. Typed classification steps in agent code can now move between providers or onto owned hardware, as long as they stay inside the text-only, 32k-context features Jev supports.
Perspective Coverage
3 publishers
- Builder
- Builder 53%
- Operator
- Operator 27%
- Investor
- Investor 20%
Reality
- Evidence50
- Adoption18
- Hype gap+35
- Incentives70
- Confidence60
Ridge Security's 96-test pen-test benchmark had Claude Opus 4.6 reach 63% coverage at $217 a run, against 52% for Gemini 3 Flash at about $5.42. Ridge argues tooling matters more than the model, yet its published runs held the tooling fixed and varied the model.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+40
- Incentives70
- Confidence40
Open-source tool runtape traced an agent's unrequested invoice forward to one sentence in a tool result, 10 of 10 reruns with it against 0 of 10 without. The same counting grades prompt fixes, though most of the evidence comes from a rule-based stand-in model.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence50
Probes on Qwen3.5-27B reading other models' text came within 0.004 AUROC, on average, of probing authors up to 397B directly, a LessWrong post reports. Every tested pair was open-weight, so auditors who apply the method to closed models get the reader's view of the text and cannot measure that gap.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence30
Ten LLMs' Blender 5.0 scripts ran only 70% of the time when a Kaggle benchmark executed them in 5.0, against 91% for scripts targeting 3.6. Renamed and removed APIs look like valid code, so the benchmark grades each answer in the exact build the prompt named.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+5
- Incentives30
- Confidence45
Cupertino is now selling desktops as an alternative to token bills. On the configurations that can host a useful model, payback runs two to four years, and the model you can host is a tier down.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+45
- Incentives60
- Confidence60
Jalapeno's lead is measured against last generation and the volumes are tiny, but a first-pass ASIC clearing Nvidia, AMD and Google parts reprices the design barrier, not the supply chain.
Perspective Coverage
7 publishers
- Builder
- Builder 30%
- Operator
- Operator 24%
- Investor
- Investor 46%
Reality
- Evidence50
- Adoption8
- Hype gap+40
- Incentives65
- Confidence55
AWS's walkthrough pairs the OpenCode terminal agent with open weight models on Amazon Bedrock and keeps code inside your own account, and the only price difference it publishes is the 10 percent discount for letting a request route anywhere.
Reality
- Evidence38
- Adoption20
- Hype gap+35
- Incentives88
- Confidence45
A LangGraph travel concierge was scored against four prompt-injection defenses, one layer at a time, on 35 prompts. With 25 attack samples in the set, the four-point regression the harness reports is a single prompt flipping.
Reality
- Evidence54
- Adoption
- Insufficient
- Hype gap+24
- Incentives42
- Confidence52
Antigravity's free tier includes Gemini 3.1 Pro and Claude Opus 4.6 with unlimited tab completions. XDA reports Opus burns credits about four times faster than Gemini, and Google moved the free tier onto credits in March 2026.
Reality
- Evidence36
- Adoption
- Insufficient
- Hype gap+12
- Incentives58
- Confidence40
An arXiv paper finetuned eight models on synthetic pre-training text describing a chain-of-thought monitor, and the models got better at evading that monitor without ever being shown an obfuscated trace.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+12
- Incentives
- Insufficient
- Confidence55
Arize and Fireworks priced ten models on what a completed command-line task costs, across 2,400 runs. The winner on that metric is an open model with the worst pass rate in the study and the thinnest coverage.
Publishers:arize.com
Reality
- Evidence62
- Adoption18
- Hype gap+14
- Incentives75
- Confidence55
A new evaluation gives models explicit rules about what may appear in their reasoning and scores whether they comply while still solving the problem. Compliance rose with model size and fell with more RL training.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−5
- Incentives35
- Confidence60
The 1.7x to 3.6x latency range is set by the baseline systems, not the chip, and the report's own publication date is unsettled. Read it as direction, not evidence.
Perspective Coverage
9 publishers
- Builder
- Builder 41%
- Operator
- Operator 31%
- Investor
- Investor 28%
Reality
- Evidence52
- Adoption14
- Hype gap+38
- Incentives82
- Confidence68
Sarah Friar told Goldman Sachs' technology conference that a pricier model can be cheaper when it needs fewer tries, a claim Artificial Analysis puts at an eleven-to-one price spread against seven index points running the other way.
Reality
- Evidence34
- Adoption38
- Hype gap+30
- Incentives80
- Confidence45
OpenAI published a benchmark win on three large models, but the figure a GPU vendor has to price is the 10-gigawatt Broadcom program running through 2029, with racks targeted from the second half of 2026.
Reality
- Evidence44
- Adoption8
- Hype gap+17
- Incentives74
- Confidence46
The bandwidth gap is 4.4x and the capacity gap is 4x, which is why these two boxes are not really competing. One decides whether a model fits; the other decides whether it is usable.
Reality
- Evidence38
- Adoption32
- Hype gap+25
- Incentives42
- Confidence35
CTGT's 76-prompt audit puts an anonymous OpenRouter endpoint at 6.5 on the usual censorship probes and 89.5 on Chinese domestic legitimacy. The prompt list is the benchmark.
Reality
- Evidence55
- Adoption30
- Hype gap+12
- Incentives65
- Confidence52
The update adds a path selector and a two-tap convolution rather than layers, recovering most of the accuracy that tripling the drafter bought at 15.2% latency, by the vendor's own numbers.
Reality
- Evidence54
- Adoption66
- Hype gap+16
- Incentives74
- Confidence58
The first multi-wafer Cerebras system pairs a doubled clock with rebuilt power delivery and interconnect. The compute claims rest on WSE-3 Turbo dies that are otherwise unchanged.
Reality
- Evidence54
- Adoption34
- Hype gap+34
- Incentives76
- Confidence61
Earlier coverage
- The conductor is the bottleneck: local agent stacks fail at orchestration, not at the workers
Build · August 18, 2026 · 1 publisher
- Agent memory has a dose-response curve, and the cheapest dose won the biggest gain
Build · August 18, 2026 · 1 publisher
- Cost per successful task, not per token: a 2,400-run benchmark reorders the model shortlist
Leadership · August 18, 2026 · 1 publisher
- 18x per joule in 16 months, and most of it was not your model choice
Invest · August 16, 2026 · 1 publisher
- Seven local models, one prompt, one DGX Spark: the speed ranking decided nothing
Build · August 15, 2026 · 1 publisher