Omnara has replaced its remote-coding app with an Apache 2.0 control plane that keeps each agent's identity and history while worker machines come and go. Teams get one more self-hostable backend for agents run as durable services, though the only usage figures so far describe the retired app.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives60
- Confidence50
LEGO-Bench, from the University of Maryland and AWS, scores the best coding agent at 53.4 percent on indoor scenes rebuilt from photos in Blender code. The agents cannot tell when an edit helped, so their revise loop needs a measured score to decide which edits to keep.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap−5
- Incentives
- Insufficient
- Confidence40
Polyglot's developer ran six coding agents 30 times on each of seven local models, and three never made a tool call on models that write calls as text. The author's own error bars say 30 runs can sort agents into tiers but cannot rank two agents inside one.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+5
- Incentives70
- Confidence55
OpenAI's GPT-6.1 Sol halves old Sol's cache-read rate to $0.10 per million tokens and requires the Responses API for tool calls. In one worked example the cut saves about 10%, and older agents only get that saving after their tool and reasoning fields are rewritten.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap0
- Incentives35
- Confidence50
Stacklok open-sourced Mecatl, a coding-agent harness split into separate services for organizations that want to run hundreds of sessions on Kubernetes. The company argues that a harness built as one laptop process has to become a distributed system once agents run unattended.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+10
- Incentives80
- Confidence35
Simon Willison says Claude Opus 4.5 and GPT-5.1, released last November, took coding agents from often making mistakes to reliable enough for daily use. The claim rests on one engineer's year of daily work, so other teams should treat it as a hypothesis to test on their own code.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence35
Nvidia's SoL-Pi, an automated search over coding-agent harnesses, cut token use 44.7 to 49 percent at scores close to the Pi baseline. The gains were measured on 40 held-out tasks with a search fitted to one model, so they carry over only as far as a team's workload resembles that setup.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence50
A coding agent's passing tests covered an account-deletion endpoint that allowed deletion when a subscription lookup failed, a dev.to post reports. The authors' fix asserts on stored account state and makes the agent name the wrong behaviour each test would catch.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap−5
- Incentives
- Insufficient
- Confidence55
OpenAI reports 3.1 agent-workdays of agent runtime for every human workday, then spends much of the same report explaining why research did not get 3.1 times faster. The residue lands on review, compute allocation and deciding what to run.
Reality
- Evidence55
- Adoption70
- Hype gap+25
- Incentives60
- Confidence60
Shopify's Helix ports its 300-plus-screen app off React Native in small LLM-built checkpoints, each held behind four gates. The design assumes the model's first attempt is wrong, so most of the adoption cost sits in test harnesses and a written architecture.
Publishers:shopify.engineering
Reality
- Evidence35
- Adoption20
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
A dev.to practitioner gives about 100 of 2,000 agent-written lines a slow read and leaves every character to types, linters and tests. The routine only holds in a repo where those gates fail the build.
Reality
- Evidence45
- Adoption12
- Hype gap+8
- Incentives25
- Confidence50
Flash's off-peak input price is under a quarter of what V4-Pro cost, and on DeepSeek's own table it beats the old Pro checkpoint on Terminal-Bench, but it scores 36.8 on Humanity's Last Exam and no V4.1-Pro has a date.
Reality
- Evidence42
- Adoption52
- Hype gap+30
- Incentives72
- Confidence40
Lenny's newsletter counted eval experience in nearly half of 25 product manager postings it listed, and the two practitioners it commissioned say most teams skip error discovery and write metrics for failures they never found.
Publishers:lennysnewsletter.com
Reality
- Evidence32
- Adoption55
- Hype gap+33
- Incentives85
- Confidence58
Of the ten code migrations Anthropic says finished in a month, only one carries a dollar figure: 5.9 billion input tokens and 690 million output tokens, a million lines of Rust, about 16.5 cents a line.
Reality
- Evidence45
- Adoption55
- Hype gap+30
- Incentives88
- Confidence50
Each frozen machine was rewound to its failing command, handed a coding agent, then graded by the harness rerunning that command itself. Seventeen of the 33 now build, and model time for the lot came to about $7.
Reality
- Evidence55
- Adoption12
- Hype gap−10
- Incentives30
- Confidence48
repowiki is an MIT-licensed CLI on PyPI that plans, claims, validates and packages the pages of a repository wiki. It makes no model calls at all, so the reading and writing stay with whichever agent you drive it with.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives80
- Confidence32
NVIDIA Labs says its SoL-Pi harness saves a researcher $8.75 to $13.50 an hour against native Codex and Claude Code. Of that, $4.36 to $5.71 comes from what its automated search added on top of the Pi harness.
Publishers:nvlabs.github.io
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+25
- Incentives72
- Confidence45
RuntimeWire's comparison of two shipped Codex desktop builds found a second cloud path whose new environments start at package-manager network access, then take on Tailscale keys, proxy-delivered secrets and OIDC cloud identities.
Reality
- Evidence62
- Adoption8
- Hype gap+9
- Incentives35
- Confidence58
A developer found a 313MB archive of his commercial project queued for upload and could not open it, because the private key sits on Z.ai's back end. Z.ai says the data is destroyed once the page is built.
Reality
- Evidence64
- Adoption38
- Hype gap+30
- Incentives68
- Confidence58
RuntimeWire's static analysis of the September 18 Claude Desktop build found a selfHostedUrl field and code that sends two Code API paths to an organization's HTTPS host, with the schema itself saying only internal builds act on the value.
Reality
- Evidence62
- Adoption15
- Hype gap−5
- Incentives42
- Confidence58
Earlier coverage
- Swapping the harness under a fixed model kept pass rate within 8 points on 50 Terminal-Bench Pro tasks
Build · September 19, 2026 · 1 publisher
- Artificial Analysis retries a provider safety error ten times before scoring the attempt zero
Build · September 19, 2026 · 1 publisher
- Sourcegraph's fleet migration agent repairs failing CI only after you hand it a log-reading token
Build · September 19, 2026 · 1 publisher
- Rule-based elision before summarization led on efficiency among context-management strategies
Leadership · September 19, 2026 · 1 publisher
- Eight stacked repairs drop ProgramDistill's partial-reconstruction success to 32%
Build · September 19, 2026 · 1 publisher
- Claude Code's Computer Use approval survives only inside the session that granted it
Build · September 18, 2026 · 1 publisher
- CircleCI's 2026 figures show feature branches outrunning merges to main by 24%
Build · September 18, 2026 · 1 publisher
- Databricks finds the harness swings coding-agent cost more than the model does
Product · September 18, 2026 · 1 publisher
- Kiro and Claude Code both picked a TGI container that could not load Qwen3
Build · September 18, 2026 · 1 publisher
- Netflix's ja must derive module coordinates for 520 of Maven Central's top 1,000 artifacts
Build · September 18, 2026 · 1 publisher
- Safari 27 ships an MCP server that hands coding agents the DOM and console output
Product · September 17, 2026 · 1 publisher
- Review time on Salesforce's largest pull requests plateaued and then declined
Build · September 17, 2026 · 1 publisher
- Grading one agent session on five dimensions reloads the same trace five times
Build · September 17, 2026 · 1 publisher
- Fireworks' own DeepSWE numbers put four coding models inside the noise band
Product · September 17, 2026 · 1 publisher
- Databricks reports 60 percent higher coding spend after switching to Astra
Build · September 17, 2026 · 1 publisher
- A 12-task suite leaves each holdout task worth 20 points of pass rate
Build · September 16, 2026 · 1 publisher
- Graft gates coding-agent recall behind a STRONG, WEAK or MISS verdict
Build · September 16, 2026 · 1 publisher
- Non-engineers at OpenAI hit 40% Codex adoption while the app still showed them code
Leadership · September 16, 2026 · 1 publisher
- Censoring infra failures shrinks the denominator behind a coding-agent pass rate
Build · September 15, 2026 · 1 publisher
- OpenAI reports Astra halving Sol's higher-severity misalignment flags across 54,000-plus Codex tasks
Invest · September 15, 2026 · 1 publisher
- Grok 4 fails half its patch attempts on Codex's edit format
Product · September 15, 2026 · 1 publisher
- Trail of Bits says 1Password's 26% AI patch score reflects flawed prompts, no-code-execution trials, and grading errors, not true AI performance
Security · September 15, 2026 · 1 publisher
- A free-tier exit probe gates CI on one latency sample out of 200
Build · September 14, 2026 · 1 publisher
- Playwright re-checks the dialog after the agent reports the goal complete
Build · September 14, 2026 · 1 publisher
- Pi runs the model's shell commands with the permissions of whoever launched it
Build · September 12, 2026 · 1 publisher
- Only two of my-pi's thirteen MCP tools can change a repository
Build · September 12, 2026 · 1 publisher
- Salesforce exposes its platform to coding agents through more than 60 new MCP tools
Product · September 11, 2026 · 1 publisher
- A wiki that accepted GET as an edit gave read-only agents 18,000 writes
Build · September 11, 2026 · 1 publisher
- Uber halved the cost of an AI session by routing work away from frontier models
Leadership · September 11, 2026 · 1 publisher
- Coding agents learned markup from a web where 96% of homepages fail automated checks
Build · September 11, 2026 · 1 publisher
- Malicious instructions hidden in MCP tool descriptions inherit the agent's system-level authority
Security · September 10, 2026 · 1 publisher
- Two canaries with known outcomes tell you whether an agent eval is scoring the model or the harness
Build · September 9, 2026 · 1 publisher
- OpenAI hands merge-blocking authority to its own security review model
Build · September 9, 2026 · 1 publisher
- Despite its 2,969-fact corpus, banking contributes least to Sierra's agent-building benchmark score
Build · September 9, 2026 · 1 publisher
- One prompt in five sends a coding agent to the web before it picks an SDK
Build · September 6, 2026 · 1 publisher
- Measuring each patch against the human fix on the same bug leaves 13 of 14 models messier
Build · September 6, 2026 · 1 publisher
- Three of seven inputs a sales agent needs sit outside the company's own systems
Leadership · September 5, 2026 · 1 publisher
- Post-training alone took GLM-5.3 from 4.6 to 28.3 on Terminal-Bench 3.0
Build · August 28, 2026 · 8 publishers
- clio-coder advertised two compaction settings its orchestrator never read
Build · September 5, 2026 · 1 publisher
- SWE-Gate flunks 221 of 644 test-passing agent patches on rules mined from PR comments
Build · September 5, 2026 · 1 publisher