Microsoft researcher Jiapeng Li's Limbo benchmark found AI agents duplicated side-effecting writes in 56% and 74% of in-flight and double-delivery episodes. Offering idempotency keys cut duplicates from 28% to 4%, so tools that move money or send mail should require them.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives25
- Confidence40
OpenAI's Help Center says Memory in ChatGPT Space can write a user's personal context onto a shared page, where every viewer can read it. OpenAI's safeguard is a review before sharing, a check that runs once while collaborators' agents keep editing.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives30
- Confidence45
TypeScript 7.0, the Go-native compiler GA since July, removes the options 6.0 deprecated, so ignoreDeprecations only postponed the break. The audit has to run on 6.0, while tsc still names each option it will drop.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence55
Google, Anthropic and OpenAI shipped four model releases in 12 days, and each one's own docs list ways that code written for the earlier version now fails. A swap of the model ID is a dependency upgrade and needs contract tests at the provider boundary before it ships.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence55
SvelteKit 3.0 requires Node 22.17, TypeScript 6 and Svelte 5.56.4 or newer before any app code changes. Teams have a toolchain project to schedule and ship before they reach the config move and the rewrite of every $lib import.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence62
Transactional mail can pass SPF and DKIM and still fail DMARC when neither aligns with the visible From domain, a dev.to Node.js guide shows. Because forwarding breaks SPF, the author treats aligned DKIM as the path to protect.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence50
Pi Durable ports the Pi agent harness to TypeScript and checkpoints every step to one of three stores, so crashed agents resume where they stopped. Whether it can replace hand-built resume code depends on how it treats a tool call cut off mid-flight.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence35
Both Astra + Luna runs of a SQLite job queue cleared their acceptance tests yet computed leases from a clock read taken before the write lock. An audit exposed it by holding one writer behind a barrier and advancing an injected clock during the other's wait.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence45
Microsoft's Go port of the TypeScript 7 compiler cut VS Code's full build from 125.7 to 10.6 seconds in its benchmarks. Tools that import TypeScript as a library may still need the TypeScript 6 API, so the upgrade belongs on a throwaway branch first.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence45
Node.js marks type stripping Stable in versions 25.2 and 24.12, letting node script.ts run TypeScript without ts-node, tsx or a build step. Node only deletes types, so the switch holds for code that stays inside TypeScript's erasable subset.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−5
- Incentives
- Insufficient
- Confidence55
PEP 827 would add conditional, comprehension and member types so Python annotations can infer return types the way Prisma's TypeScript ORM does. Michael J. Sullivan spent his Language Summit slot on how those annotations get stored for runtime users like Pydantic and FastAPI.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence60
PDF libraries can write a form field's new /V value and leave the old /AP appearance on the page, a dev.to guide on redaction warns. For personal data, the guide makes a check of the rendered pixels the release gate, because exit codes pass either way.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+3
- Incentives22
- Confidence52
Timescale says Claude wrote nearly all the code and ran all 41 experiments that lifted its LoCoMo memory score from 0.392 to 0.666 F1 in six days. By the team's own account, the calls that decided the result came from a person checking whether each metric measured what it claimed to.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+5
- Incentives45
- Confidence40
Vercel Labs' experimental ScriptC compiles TypeScript to native binaries that started in 1.78ms against Node's 61.78ms in one benchmark. The gain holds for short, statically typed programs, while framework code and sustained compute both ran slower than Bun or Node.
Reality
- Evidence55
- Adoption15
- Hype gap+20
- Incentives
- Insufficient
- Confidence60
One team's in-app AI agent never saw 21 working handlers because its tool manifest drops any entry without a description string. The text has to be written beside each handler's definition, so the team now runs gap tests that fail when a feature ships blank.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence50
Publishers onboarding custom domains should confine wildcard DNS to routing and verify three mail records per tenant, a dev.to guide argues. That costs extra writes and waiting, and it gives every customer its own rollback boundary and audit trail.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence50
Same declared ESLint range, opposite outcomes: 39 jsx-a11y rules run clean on 10.9.0 while 38 of eslint-plugin-react's 101 throw, and one line of config moves 32 of them.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+8
- Incentives40
- Confidence66
Exact-substring scoring put a developer's 700-line RAG tool at a 65% retrieval hit-rate at k=3, 13 points below what a token-overlap scorer found. The bug also made extra retrieved chunks look worthless, and a 20-question test set added 15 points of noise.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence55
Temperô's iFood integration checked a field holding 'PLC' against cases written for 'PLACED', so every real event hit a silent no-op for weeks. On a poll-and-acknowledge queue, a default branch that logs nothing lets events disappear without an error.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives25
- Confidence55
Stephen Toub of Microsoft writes that AI agents wrote most of the code and a single developer drove the port in a few months while the rest of the team kept expanding the runtime. The post puts the speedup at orders of magnitude.
Reality
- Evidence45
- Adoption60
- Hype gap+35
- Incentives80
- Confidence55
Earlier coverage
- Capping fetch depth took Linear's slowest CI gate from 94 seconds to 20
Build · September 21, 2026 · 1 publisher
- Cursor's Rollouts bot grades a deploy inconclusive when the telemetry cannot decide
Build · September 24, 2026 · 1 publisher
- An eight-attempt history has to climb 3.5 points before CogniPrep draws an arrow
Build · September 24, 2026 · 1 publisher
- Staying lazy until the sort cuts a chain's 80,000-object peak to about 2,000
Build · September 23, 2026 · 1 publisher
- Dev's math suggests the supervisor LLM drove most of a reported 70% token cut, if handoff truly costs zero tokens
Build · September 23, 2026 · 1 publisher
- Anthropic wants MCP agents to write code instead of loading every tool definition
Leadership · September 23, 2026 · 1 publisher
- On the plan screen, Opus opens on the set clash while Astra and Sol keep it below a repeated hero
Build · September 22, 2026 · 1 publisher
- Salesforce engineers call flat review time on the biggest pull requests a sign of disengagement
Build · September 22, 2026 · 1 publisher
- Bun's million-line Rust port ran $165,000 through Anthropic's API meter
Invest · September 22, 2026 · 1 publisher
- Handing a coding agent the frozen machine fixed 17 of 33 broken hackathon builds
Build · September 22, 2026 · 1 publisher
- Bun's 535,496-line Zig codebase was rewritten in Rust and validated against TypeScript tests
Build · September 22, 2026 · 1 publisher
- Node 24 overwrites a 34-character type annotation with 34 spaces to keep columns exact
Build · September 22, 2026 · 1 publisher
- In one worked example, a coding agent waits 93 percent of its cycle on CI
Build · September 22, 2026 · 1 publisher
- Generated flex rows need min-w-0 before Tailwind's truncate will clip a long username
Build · September 21, 2026 · 1 publisher
- Guest play was deleted not because of the signed permit, but because it reopened anonymous access to authored content for an unmeasured funnel
Build · September 21, 2026 · 1 publisher
- Atlassian ships @forge/bridge with view.submit() typed to accept any object
Build · September 20, 2026 · 1 publisher
- AI agents ported Bun's 535,496 lines of Zig to Rust, validated by a language-independent test suite
Build · September 20, 2026 · 1 publisher
- Month two of an unreviewed AI test pipeline left 30 percent of generated Playwright tests flaky
Build · September 20, 2026 · 1 publisher
- Microsoft puts Rust on its top internal support tier alongside C++, C# and TypeScript
Leadership · September 20, 2026 · 1 publisher
- Lootr Tools refuses a drop rate that arrives without a provenance status
Build · September 19, 2026 · 1 publisher
- Reactive Agents repairs the almost-right tool call so the run keeps going
Build · September 19, 2026 · 1 publisher
- A log dashboard on default credentials hands over every plaintext password
Build · September 19, 2026 · 1 publisher
- Next.js 16 asks a glibc 2.28 host for GLIBC_2.29 before it reads next.config.ts
Build · September 19, 2026 · 1 publisher
- A one-developer UI library rebuilt the template type-checking Angular ships by default
Build · September 19, 2026 · 1 publisher
- A full arcade redesign stopped at eight logic.ts files
Build · September 19, 2026 · 1 publisher
- A recovery-decision event can be written before the provider's own event id exists
Build · September 18, 2026 · 1 publisher
- A shadcn linter tells coding agents which fix the design system allows
Leadership · September 18, 2026 · 1 publisher
- An execution boundary lets the application refuse an agent's proposed tool call
Build · September 18, 2026 · 1 publisher
- Real-SWE licenses private production codebases to score coding agents on real business tasks
Build · September 18, 2026 · 1 publisher
- AsyncLocalStorage moves the missing-tenant bug from the call site to the request boundary
Build · September 18, 2026 · 1 publisher
- Freeze the pagination contract before the agent writes the list handler
Build · September 18, 2026 · 1 publisher
- Six runtime events map to six of eight UI states in this Angular agent walkthrough
Build · September 17, 2026 · 1 publisher
- ArchUnitTS fails an architecture rule that inspected zero files
Build · September 17, 2026 · 1 publisher
- Don Watch replays the 2019 Sheffield flood from gauge records inside one browser tab
Build · September 17, 2026 · 1 publisher
- Pre-assigning a request ID turns a timed-out generation call into a resumable attempt
Build · September 17, 2026 · 1 publisher
- Reinjecting the bug is how this experiment scored its coding agent's patches
Build · September 17, 2026 · 1 publisher
- 40,000 inventory updates outlive the 900-second cron runner that retries them
Build · September 16, 2026 · 1 publisher
- A Cloudflare Worker generates the code_challenge IFTTT's redirect leaves out
Build · September 16, 2026 · 1 publisher
- Sync pill flips offline on an explicit OS false or a lost PowerSync connection
Build · September 16, 2026 · 1 publisher
- Flipping the default to 402 got this Korean tax-data MCP server into the x402 Bazaar
Build · September 15, 2026 · 1 publisher