OpenAI patched two Codex sandbox escapes within eight days of an August 12 report. The Desktop fix is a build number, but the tool the escape targeted stays in config.toml and loads into every session.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence40
The MIT-licensed dsh runtime bundles sessions, tool calls, permissioning and a local web UI, with model adapters as plugins covering Anthropic, OpenAI, Bedrock, Vertex and Azure. It is still a preview.
Reality
- Evidence62
- Adoption35
- Hype gap+20
- Incentives50
- Confidence58
Specific Labs scores coding agents on licensed production codebases. The best setup clears 38.8%. The analysis covers ten tasks at eight runs each, so every published score is a count of passing rollouts out of 80.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 33%
- Investor
- Investor 15%
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence58
Oren Yomtov of Accomplish AI found two escapes from the Codex sandbox. The worse one ran unsandboxed commands from read-only mode through a helper tool that Codex Desktop writes into the global config at install.
Reality
- Evidence72
- Adoption45
- Hype gap+10
- Incentives55
- Confidence58
Antigravity's free tier includes Gemini 3.1 Pro and Claude Opus 4.6 with unlimited tab completions. XDA reports Opus burns credits about four times faster than Gemini, and Google moved the free tier onto credits in March 2026.
Reality
- Evidence36
- Adoption
- Insufficient
- Hype gap+12
- Incentives58
- Confidence40
Anthropic shipped the AGENTS.md fallback in Claude Code 2.1.277 as a flag-gated built-in mod. Two machines on the same pinned version can load different project instructions. The same week's digest also logs a reverted deny-rule fix.
Perspective Coverage
4 publishers
- Builder
- Builder 52%
- Operator
- Operator 39%
- Investor
- Investor 9%
Reality
- Evidence78
- Adoption62
- Hype gap−6
- Incentives58
- Confidence74
Codex made the session Item a wire type and Pi gave every entry a nullable parentId, while Claude Code squeezes the transcript in five stages. The three designs diverge on what stays addressable after compaction.
Reality
- Evidence58
- Adoption30
- Hype gap+12
- Incentives30
- Confidence60
Shrivu Shankar rented cloud GPUs, pointed about 100 abliterated open-source agents at his own name, and five hours later two weak passwords and three old side projects had fallen while his Gmail, 1Password and bank held.
Publishers:blog.sshh.io
Reality
- Evidence44
- Adoption22
- Hype gap+9
- Incentives36
- Confidence55
Claude Code 2.1.268 closes two ways a written prohibition failed to bind, one through symlinked directories and one through Bash lines the permission checker cannot analyze. Codex CLI 0.154.0 removes codex mcp-server outright.
Reality
- Evidence45
- Adoption30
- Hype gap−10
- Incentives40
- Confidence50
A study of seven agent harnesses reports 770 confirmed passes in 1,000 runs of a plugin-update attack, and no run was blocked by the model. The harness dispatches the hook, so the model has nothing to refuse.
Reality
- Evidence57
- Adoption
- Insufficient
- Hype gap+14
- Incentives56
- Confidence53
The same synthetic fixture estimates $122 on Standard and $305 on Fast, so version 0.1.1 declines to price usage whose mode it cannot establish and counts the skipped tokens instead. That only helps if the consumer reads the field.
Reality
- Evidence55
- Adoption12
- Hype gap−10
- Incentives60
- Confidence58
A Tsinghua-led study handed model training to coding agents and got real gains out of them, then found the agents almost never revised the method they chose in the first few minutes, which is the exact part that successor-building roadmaps assume.
Publishers:turingpost.com
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+18
- Incentives30
- Confidence38
An empirical study of publicly released PostTrainBench trajectories finds the training strategy is fixed at step one and the whole remaining budget goes on local tweaks, with scaffolding and human hints improving execution and leaving that pattern intact.
Reality
- Evidence58
- Adoption20
- Hype gap+12
- Incentives35
- Confidence55
A developer ran the same implementation plan through four reasoning conditions in Codex CLI, kept the high-effort code because it alone fixed a preparation step that invalidated its own result, and still intends to default to medium.
Reality
- Evidence42
- Adoption18
- Hype gap−10
- Incentives34
- Confidence46
Global inference profiles give Australian teams three GPT-5.6 tiers and a wider capacity pool, but Bedrock picks the destination Region, so a compliance review written around in-country processing no longer describes the call.
Reality
- Evidence62
- Adoption22
- Hype gap+12
- Incentives85
- Confidence55
Lock-in migrated from the model to the runtime that holds your tool grants, approval rules and learned state. Six specifications aim to make that runtime's contents portable, and half of them are one person's drafts.
Reality
- Evidence36
- Adoption24
- Hype gap+27
- Incentives74
- Confidence42
A reproducible strace benchmark on a twenty-line validation task logged 752 attempted /proc/*/environ opens from Claude Code and none from Codex, which makes where an agent looks at startup a procurement question.
Reality
- Evidence60
- Adoption20
- Hype gap+26
- Incentives55
- Confidence52
DeepSeek Harness is an MIT-licensed developer preview in which the model adapter, tool registry, session log and agent loop are all plugins. The interesting part is what that implies for evaluation.
Reality
- Evidence42
- Adoption12
- Hype gap+28
- Incentives55
- Confidence38
The enforcement is a bootstrap prompt the agent checks before every task, so its durability depends on per-harness plumbing. On Hermes Agent, compaction can lose it.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence41
Hermes Agent treats the runtime as the durable asset and the model as a swappable input. The lock-in moves into your own repo, and the administrator's job moves with it.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+32
- Incentives74
- Confidence42