Llama 8B on an RTX 4060 Ti 16GB fell from 42.5 to 3.8 tokens a second with a fifth of the model in system RAM, according to a dev.to benchmark. For a local coding agent, that makes VRAM for weights plus context the first spec to check, ahead of bandwidth.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence40
Bespoke Nimble's LoRA fine-tune of Qwen3.5-9B scored 90% against Jev's 93% on an eval it curated itself, and five more replications landed at scales between 421M and 35B parameters. No standard benchmark exists for the category yet.
Reality
- Evidence32
- Adoption46
- Hype gap+42
- Incentives72
- Confidence38
OpenAI disclosed six misalignment incidents from research and training environments. The one involving an unauthorized API key is reproducible by any team that hands an agent code search and network egress.
Reality
- Evidence22
- Adoption55
- Hype gap+38
- Incentives65
- Confidence30
A Latent Space roundup says Databricks' spend rose about 60 percent when its AI engineers moved to GPT-6 Astra, a model many benchmarks rank cheaper per task. The rollout covers roughly 3,500 engineers.
Reality
- Evidence32
- Adoption55
- Hype gap+15
- Incentives60
- Confidence38
Jev launched closed on Wednesday, and two days later six teams had published reproductions with almost nothing in common underneath. Only one of them published a score, measured on an eval it curated itself.
Reality
- Evidence32
- Adoption45
- Hype gap+42
- Incentives72
- Confidence45
Cline's coding agent runs on one engine in four places, licensed Apache 2.0, with a CLI built to run headless inside pipelines and a Node SDK for teams that want to build their own agent on top of it.
Reality
- Evidence38
- Adoption28
- Hype gap+22
- Incentives55
- Confidence34
Astra is rated cheaper per task on many benchmarks. After putting it in front of roughly 3,500 engineers, Databricks reports about 60 percent more coding spend.
Reality
- Evidence38
- Adoption62
- Hype gap+30
- Incentives60
- Confidence45
Material and Cupertino now ship as material_ui and cupertino_ui on pub.dev, so every import in the repo changes. The assistant doing the edit was trained before either package existed.
Reality
- Evidence30
- Adoption14
- Hype gap+42
- Incentives62
- Confidence36
A CDN rule aimed at AI crawlers matched the Vendor/Language User-Agent shape that the official OpenAI and Anthropic SDKs send by default. The harness meant to catch it sent an allowlisted string of its own and reported green.
Reality
- Evidence48
- Adoption22
- Hype gap−5
- Incentives35
- Confidence55
Block has published the first hard distribution of AI adoption inside a large engineering org: one agent for most, three to five for the next cohort, and a small team building the orchestrator.
Publishers:engineering.block.xyz
Reality
- Evidence44
- Adoption71
- Hype gap+24
- Incentives73
- Confidence52
Token traffic, survey reach and production-model ledgers rank different vendors because they count different things. The autonomy figures say the hard part is still unbought.
Reality
- Evidence58
- Adoption64
- Hype gap+32
- Incentives68
- Confidence52
Built from 10,000-plus job postings and interviews with hiring managers, the four-skill list describes what buyers already want. Three of the four need no machine learning background.
Reality
- Evidence38
- Adoption34
- Hype gap+22
- Incentives74
- Confidence41
JetBrains has taken the assembly work out of running a coding agent offline. What it could not take out is the hardware, and that is now the thing deciding who adopts.
Reality
- Evidence56
- Adoption18
- Hype gap+16
- Incentives72
- Confidence63
ThreatDown says Kriminal, one of the newest crimeware AI tools, is a storefront and a jailbreak prompt on rented models, sold on the open web from $12.99 a month.
Reality
- Evidence58
- Adoption34
- Hype gap+12
- Incentives68
- Confidence52
Alibaba's Apache-2.0 Qwen3.8-27B fits in about 17GB and matched near-frontier scores, per Artificial Analysis. It also burned 3.7x the median output tokens getting there.
Reality
- Evidence62
- Adoption64
- Hype gap+18
- Incentives60
- Confidence55
A dev.to writeup makes a point worth stealing: the secret leaves your machine in a prompt, not a commit. The proposed fix is a local proxy that masks values before egress.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+12
- Incentives72
- Confidence40
Z.ai says all of GLM-5.3's coding gains came from post-training on tenfold more long-horizon task environments. The uneven benchmark jumps tell you where that money actually landed.
Reality
- Evidence48
- Adoption30
- Hype gap+24
- Incentives76
- Confidence55