Skip to content

Topic

Coding Agents

AI software agents that autonomously read, write, and refactor code in real repositories, often invoking tools or shell commands to complete tasks.

Current stories

build1 publisher

Simon Willison credits two November model releases with making coding agents reliable for daily use

Simon Willison says Claude Opus 4.5 and GPT-5.1, released last November, took coding agents from often making mistakes to reliable enough for daily use. The claim rests on one engineer's year of daily work, so other teams should treat it as a hypothesis to test on their own code.

Reality

Evidence30
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence35
build1 publisher

Shopify's Helix has LLMs port a 300-plus-screen app one reviewed checkpoint at a time

Shopify's Helix ports its 300-plus-screen app off React Native in small LLM-built checkpoints, each held behind four gates. The design assumes the model's first attempt is wrong, so most of the adoption cost sits in test harnesses and a written architecture.

Publishers:shopify.engineering

Reality

Evidence35
Adoption20
Hype gap+20
Incentives
Insufficient
Confidence40
invest1 publisher

Nearly half of one newsletter's 25 product openings ask for eval experience

Lenny's newsletter counted eval experience in nearly half of 25 product manager postings it listed, and the two practitioners it commissioned say most teams skip error discovery and write metrics for failures they never found.

Publishers:lennysnewsletter.com

Reality

Evidence32
Adoption55
Hype gap+33
Incentives85
Confidence58

Earlier coverage

  1. Swapping the harness under a fixed model kept pass rate within 8 points on 50 Terminal-Bench Pro tasks

    Build · September 19, 2026 · 1 publisher

  2. Artificial Analysis retries a provider safety error ten times before scoring the attempt zero

    Build · September 19, 2026 · 1 publisher

  3. Sourcegraph's fleet migration agent repairs failing CI only after you hand it a log-reading token

    Build · September 19, 2026 · 1 publisher

  4. Rule-based elision before summarization led on efficiency among context-management strategies

    Leadership · September 19, 2026 · 1 publisher

  5. Eight stacked repairs drop ProgramDistill's partial-reconstruction success to 32%

    Build · September 19, 2026 · 1 publisher

  6. Claude Code's Computer Use approval survives only inside the session that granted it

    Build · September 18, 2026 · 1 publisher

  7. CircleCI's 2026 figures show feature branches outrunning merges to main by 24%

    Build · September 18, 2026 · 1 publisher

  8. Databricks finds the harness swings coding-agent cost more than the model does

    Product · September 18, 2026 · 1 publisher

  9. Kiro and Claude Code both picked a TGI container that could not load Qwen3

    Build · September 18, 2026 · 1 publisher

  10. Netflix's ja must derive module coordinates for 520 of Maven Central's top 1,000 artifacts

    Build · September 18, 2026 · 1 publisher

  11. Safari 27 ships an MCP server that hands coding agents the DOM and console output

    Product · September 17, 2026 · 1 publisher

  12. Review time on Salesforce's largest pull requests plateaued and then declined

    Build · September 17, 2026 · 1 publisher

  13. Grading one agent session on five dimensions reloads the same trace five times

    Build · September 17, 2026 · 1 publisher

  14. Fireworks' own DeepSWE numbers put four coding models inside the noise band

    Product · September 17, 2026 · 1 publisher

  15. Databricks reports 60 percent higher coding spend after switching to Astra

    Build · September 17, 2026 · 1 publisher

  16. A 12-task suite leaves each holdout task worth 20 points of pass rate

    Build · September 16, 2026 · 1 publisher

  17. Graft gates coding-agent recall behind a STRONG, WEAK or MISS verdict

    Build · September 16, 2026 · 1 publisher

  18. Non-engineers at OpenAI hit 40% Codex adoption while the app still showed them code

    Leadership · September 16, 2026 · 1 publisher

  19. Censoring infra failures shrinks the denominator behind a coding-agent pass rate

    Build · September 15, 2026 · 1 publisher

  20. OpenAI reports Astra halving Sol's higher-severity misalignment flags across 54,000-plus Codex tasks

    Invest · September 15, 2026 · 1 publisher

  21. Grok 4 fails half its patch attempts on Codex's edit format

    Product · September 15, 2026 · 1 publisher

  22. Trail of Bits says 1Password's 26% AI patch score reflects flawed prompts, no-code-execution trials, and grading errors, not true AI performance

    Security · September 15, 2026 · 1 publisher

  23. A free-tier exit probe gates CI on one latency sample out of 200

    Build · September 14, 2026 · 1 publisher

  24. Playwright re-checks the dialog after the agent reports the goal complete

    Build · September 14, 2026 · 1 publisher

  25. Pi runs the model's shell commands with the permissions of whoever launched it

    Build · September 12, 2026 · 1 publisher

  26. Only two of my-pi's thirteen MCP tools can change a repository

    Build · September 12, 2026 · 1 publisher

  27. Salesforce exposes its platform to coding agents through more than 60 new MCP tools

    Product · September 11, 2026 · 1 publisher

  28. A wiki that accepted GET as an edit gave read-only agents 18,000 writes

    Build · September 11, 2026 · 1 publisher

  29. Uber halved the cost of an AI session by routing work away from frontier models

    Leadership · September 11, 2026 · 1 publisher

  30. Coding agents learned markup from a web where 96% of homepages fail automated checks

    Build · September 11, 2026 · 1 publisher

  31. Malicious instructions hidden in MCP tool descriptions inherit the agent's system-level authority

    Security · September 10, 2026 · 1 publisher

  32. Two canaries with known outcomes tell you whether an agent eval is scoring the model or the harness

    Build · September 9, 2026 · 1 publisher

  33. OpenAI hands merge-blocking authority to its own security review model

    Build · September 9, 2026 · 1 publisher

  34. Despite its 2,969-fact corpus, banking contributes least to Sierra's agent-building benchmark score

    Build · September 9, 2026 · 1 publisher

  35. One prompt in five sends a coding agent to the web before it picks an SDK

    Build · September 6, 2026 · 1 publisher

  36. Measuring each patch against the human fix on the same bug leaves 13 of 14 models messier

    Build · September 6, 2026 · 1 publisher

  37. Three of seven inputs a sales agent needs sit outside the company's own systems

    Leadership · September 5, 2026 · 1 publisher

  38. Post-training alone took GLM-5.3 from 4.6 to 28.3 on Terminal-Bench 3.0

    Build · August 28, 2026 · 8 publishers

  39. clio-coder advertised two compaction settings its orchestrator never read

    Build · September 5, 2026 · 1 publisher

  40. SWE-Gate flunks 221 of 644 test-passing agent patches on rules mined from PR comments

    Build · September 5, 2026 · 1 publisher