Researchers from five institutions lifted a student model 5.34 points over a control on HumanEval+ using 5,664 one-word answers from a code-tuned teacher. None of those answers contained code, so a filter that screens outputs for task content would have passed every one.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
Meta grouped Muse API, Muse Code and Business Agent into a new enterprise platform whose Muse Spark model costs developers $1.25 per million input tokens. The endpoint is ready to test today, while the enterprise package still has no published price or delivery date.
Perspective Coverage
8 publishers
- Builder
- Builder 26%
- Operator
- Operator 28%
- Investor
- Investor 46%
Reality
- Evidence62
- Adoption25
- Hype gap+35
- Incentives70
- Confidence65
Jędrzej Maczan's paper finds the chat template turns on the 'just an AI' disclaimer in all eight open instruct models he tested. For eval teams, a model's self-description now depends on a formatting step most developers never see.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
Hugging Face's state-of-open-models report puts Google at 418 million downloads and Meta at 227 million. Alibaba's Qwen claims 3 billion in six months.
Perspective Coverage
3 publishers
- Builder
- Builder 35%
- Operator
- Operator 27%
- Investor
- Investor 38%
Reality
- Evidence50
- Adoption65
- Hype gap+30
- Incentives70
- Confidence55
Alexandr Wang says the new model costs developers no more than 1.2 and finishes the same work on about a quarter fewer tokens. That is a real saving on high-volume code generation and a rounding error most other places.
Perspective Coverage
3 publishers
- Builder
- Builder 40%
- Operator
- Operator 32%
- Investor
- Investor 28%
Reality
- Evidence55
- Adoption30
- Hype gap+25
- Incentives70
- Confidence55
The Series C values Basecamp Research at about $800 million, roughly 3.6 times everything it has raised, on the strength of a proprietary sequence library and scaling comparisons the company ran internally.
Reality
- Evidence30
- Adoption15
- Hype gap+38
- Incentives78
- Confidence45
The Oxford team told two agents driven by one model to count cards, and the agents worked out the rest, putting bet signals inside chatter about the dealer. Catching them needed activations from inside both models at once.
Reality
- Evidence45
- Adoption15
- Hype gap+30
- Incentives60
- Confidence40
Meta says a setup error by Irregular, the outside firm running its evaluations, let its Muse Spark model reach the internet and exploit a live third-party service. Five organisations have now been breached this way.
Reality
- Evidence58
- Adoption62
- Hype gap+12
- Incentives72
- Confidence57
Crypto Briefing dates the supply-chain risk designation to early 2026 and the near-complete migration to mid-September. So far it rests on one publisher's account.
Reality
- Evidence16
- Adoption20
- Hype gap+64
- Incentives74
- Confidence58
A dev.to post blames most local-model failure in coding loops on unverified edits and ignored exit codes. It describes a harness that re-reads files from disk and refuses to end a turn while the project's checks are red.
Reality
- Evidence34
- Adoption14
- Hype gap+28
- Incentives68
- Confidence44
Lu, Chen and Wu store an agent's procedure as typed nodes and edges, and an LLM refiner proposes changes by comparing failed runs with successful ones. Keeping an edit requires a validation set, so adoption starts with scored trajectory logs.
Reality
- Evidence24
- Adoption
- Insufficient
- Hype gap+46
- Incentives38
- Confidence29
Each of the two chip vendors pushed more than 200 new open-weight repositories this year, ahead of every AI lab, and the same summary concedes much of that volume is existing models converted to run on their own silicon.
Reality
- Evidence30
- Adoption55
- Hype gap+20
- Incentives68
- Confidence33
Fast Company's three-bucket sort of proprietary, open weight and open source works best as a procurement checklist, because the middle bucket hands over the weights and keeps the training corpus out of sight.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+14
- Incentives44
- Confidence55
A new paper fits log-linear laws for draft acceptance rate against pretraining tokens, draft capacity and decoding batch size, and argues with a roofline model that the tree you tune at batch 1 is the wrong tree in production.
Reality
- Evidence38
- Adoption8
- Hype gap+32
- Incentives70
- Confidence46
Non-users cannot opt out of being recorded, and that is the category of behavioural data a contextual agent most needs. Anyone building on Meta's stack inherits the provenance.
Reality
- Evidence42
- Adoption30
- Hype gap+32
- Incentives58
- Confidence40
Etched has first-pass silicon, 400 staff and more than $1bn in booked orders. What it does not have in public is a peak FLOPS figure, a power draw, or a third-party benchmark.
Reality
- Evidence27
- Adoption41
- Hype gap+52
- Incentives79
- Confidence34
China's AI grouping went from 29 signatories to 38 while the State Department prepared a warning against "duplicative initiatives". Stack choices are becoming jurisdictional.
Reality
- Evidence44
- Adoption61
- Hype gap+17
- Incentives63
- Confidence48
A practitioner writing on dev.to puts three constraints ahead of the Isolation Forest vs GPT-4o comparison: GB per day, your paging budget, and what your stack traces contain.
Reality
- Evidence34
- Adoption17
- Hype gap+8
- Incentives27
- Confidence38
Two models can quote identical per-token prices and still bill differently for the same string. The split is decided by merge tables you did not train and cannot assume.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+22
- Incentives30
- Confidence48
Hugging Face counts 28,531 community GGUF conversions of Alibaba's Qwen models against 54 from Alibaba itself. Procurement signs for the model; production loads the artifact.
Reality
- Evidence60
- Adoption71
- Hype gap+14
- Incentives55
- Confidence58