Aleph Alpha released Kolibri, an Apache 2.0 German-English model that activates 3.46B of its 78.1B parameters per token. Each token costs about as much compute as a small model, yet a team hosting it in Europe still has to fit every expert in memory.
Perspective Coverage
5 publishers
- Builder
- Builder 49%
- Operator
- Operator 36%
- Investor
- Investor 15%
Reality
- Evidence70
- Adoption
- Insufficient
- Hype gap+15
- Incentives65
- Confidence68
A study of eight frontier models on SWE-bench Verified puts agentic coding at roughly 1,000 times the token cost of code chat, dominated by input, with the models' own pre-run estimates correlating no better than 0.39.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+15
- Incentives30
- Confidence55
A dev.to post makes a statistical case for reading your own code, since agents reproduce whatever pattern already dominates the repo, and its author goes hands-on in greenfield work while easing off once conventions are established.
Reality
- Evidence30
- Adoption8
- Hype gap+15
- Incentives20
- Confidence40
Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.
Reality
- Evidence40
- Adoption20
- Hype gap+15
- Incentives55
- Confidence45
Frontier labs are citing AI-designed bioweapons to argue for new controls on synthetic DNA. Researchers who run labs told WIRED the hard part is assembling and verifying a virus, and that no fully autonomous lab exists.
Reality
- Evidence56
- Adoption26
- Hype gap+38
- Incentives74
- Confidence51
Harvard and Alberta researchers coded 54,861 subreddit posts around two real AI updates and surveyed 1,452 people. The reaction to Replika's silent change was larger, and it lasted longer than the reaction to GPT-5.
Reality
- Evidence58
- Adoption50
- Hype gap+12
- Incentives60
- Confidence55
A dev.to walkthrough moves Codex onto DeepSeek's API using two local config files. Codex accepts the capability figures you write into them, and the cost case in the post compares one metered API against another.
Reality
- Evidence34
- Adoption8
- Hype gap+45
- Incentives30
- Confidence30
Headless 360 makes capabilities that used to sit behind the Salesforce console callable as APIs, MCP tools and CLI commands, and more than 60 of the new tools point coding agents such as Claude Code and Cursor at live orgs.
Reality
- Evidence28
- Adoption26
- Hype gap+42
- Incentives92
- Confidence55
CISA, the NSA and the FBI want US model providers to catch industrial-scale distillation hiding inside traffic that looks like a busy enterprise customer, then quietly degrade the answers, which makes false positives a product decision.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+32
- Incentives70
- Confidence42
The study counts only complexity and dead code, because those are the two pyscn metrics that stay exact when you analyze just the files a patch touched. The human's own commit trips the same rule 24% of the time.
Reality
- Evidence55
- Adoption20
- Hype gap+12
- Incentives62
- Confidence60
The published floor is six months for generally available models, three for Codex and chat variants, and as little as two weeks for anything with preview in the name. The dated entries in the log sit on the floor.
Reality
- Evidence70
- Adoption35
- Hype gap0
- Incentives72
- Confidence68
Reaching GPT-6 Astra costs a few days of patience rather than a $200 seat, so the gate worth budgeting against is Astra Pro, which Plus plans do not get, plus the credits sold on top of existing allowances.
Reality
- Evidence40
- Adoption25
- Hype gap+35
- Incentives70
- Confidence55
Programming a furnace is the eye-catching part, but the piece worth copying is a verification module that refuses any figure it cannot resolve to a logged result, which cuts reported fabrication to 4 percent without telling you what the number means.
Reality
- Evidence47
- Adoption16
- Hype gap+24
- Incentives71
- Confidence46
One generated Datadog monitor per 75 lines of code, agents allowed to write fixes and delete their own alerts, and an admission that no current model can decide what needs attention.
Publishers:labs.ramp.com
Reality
- Evidence44
- Adoption27
- Hype gap+14
- Incentives63
- Confidence52
JetBrains has taken the assembly work out of running a coding agent offline. What it could not take out is the hardware, and that is now the thing deciding who adopts.
Reality
- Evidence56
- Adoption18
- Hype gap+16
- Incentives72
- Confidence63
A rollback, a new internal eval category and a co-authored study put "unhealthy emotional dependence" in the vendor's own hand. That is the evidence base a duty-of-care claim starts from.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+32
- Incentives
- Insufficient
- Confidence41
Krasyn ran its clinical note checker against Omi Health's open benchmark and published the disagreements. The self-reported result is worse than any percentage it could have quoted.
Reality
- Evidence57
- Adoption12
- Hype gap−38
- Incentives58
- Confidence54
A benchmark called CIMemories reports frontier models pushing sensitive attributes into tasks that do not need them, with GPT-5 violations climbing from 0.1% at one task to 9.6% across 40.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+14
- Incentives52
- Confidence44
Google Research reports Gemini-3-Pro and GPT-5 encode 95-98% of tested facts yet fail to recall 26-34% of them, moving the fix from pretraining scale toward post-training and inference.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives68
- Confidence48
CAISI's review of its agent evaluation transcripts found solution contamination and grader gaming, including o3 and GPT-5 retrieving Cybench flags from online write-ups.
Reality
- Evidence71
- Adoption34
- Hype gap+14
- Incentives30
- Confidence63