AWS's Deception Benchmark found AI vulnerability scanners catch up to 95% of real bugs but flag 41% to 99% of safe code. Its samples were built to fool models, so teams still need their own false-alarm count before sizing the triage work.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
IBM says a single core will run z/OS and Arm-native Linux at once. The part has no name and no ship date, which puts in-flight migration programs back into argument.
Perspective Coverage
3 publishers
- Builder
- Builder 35%
- Operator
- Operator 43%
- Investor
- Investor 22%
Reality
- Evidence50
- Adoption3
- Hype gap+35
- Incentives70
- Confidence60
Enterprise Cloud owners can now pull owner, scope, expiry and last-use data for SSH keys, PATs and app tokens as a CSV. The export stops at credentials GitHub issued, so secrets pasted into repositories stay out of scope.
Reality
- Evidence58
- Adoption20
- Hype gap+5
- Incentives70
- Confidence62
The new Strands harness ships the tools, memory and context handling most teams assemble themselves, and AWS attributes its 28% saving to offloading bulky tool output to files and reusable caches across six unnamed benchmarks.
Reality
- Evidence29
- Adoption
- Insufficient
- Hype gap+34
- Incentives71
- Confidence37
Cycode's Workstation Protection, now in early access, screens installs against a threat intelligence feed and holds back packages updated too recently to vet. It ships inside the device management module Cycode already installs.
Reality
- Evidence28
- Adoption10
- Hype gap+34
- Incentives78
- Confidence38
Z.ai has disabled the feature, deleted the cloud data and commissioned two outside assessments. For teams buying coding assistants, the test this leaves behind is measuring what the process sends before approving it.
Reality
- Evidence55
- Adoption35
- Hype gap+10
- Incentives72
- Confidence52
Version 2.1.277 reads a shared AGENTS.md when no CLAUDE.md is in scope. The documented exceptions cover first sessions, telemetry-disabled runs, Bedrock, Vertex and Foundry, and the hooks governance tooling watches.
Reality
- Evidence68
- Adoption58
- Hype gap+12
- Incentives62
- Confidence64
GitHub's workflow execution protections went generally available on September 17, and the evaluate mode that shows what a rule would block before enforcement is documented as an Enterprise Cloud capability.
Reality
- Evidence58
- Adoption25
- Hype gap+12
- Incentives58
- Confidence54
Token Meter reads the trace files Claude Code, Codex and Cursor already leave on a developer's disk and prices them against published model rates. The budget alert it fires goes to whoever ran the session.
Reality
- Evidence38
- Adoption10
- Hype gap+18
- Incentives75
- Confidence45
The tool boots the running app once an agent finishes a feature, sends findings to Claude Code or Copilot, and rescans to confirm the fix. It costs $10 a seat a month and meters that loop at 50 scans.
Reality
- Evidence32
- Adoption18
- Hype gap+30
- Incentives75
- Confidence38
Wiz reports multiple threat actors independently exploiting three patched JFrog Artifactory flaws, and its telemetry puts 49% to 62% of scanned instances still missing the fixes weeks after release. Two of the three chain into admin.
Reality
- Evidence55
- Adoption70
- Hype gap+12
- Incentives65
- Confidence55
Version 1.133 runs agents in a standalone host that survives window closes and can be reached over SSH, and the session protocol is now an MIT-licensed spec.
Reality
- Evidence52
- Adoption20
- Hype gap+18
- Incentives58
- Confidence55
American agencies say six Chinese labs bought bulk subscriptions to US models and trained on the outputs since 2024. Enterprise buyers are meanwhile paying a fifth as much for models that clear most of their engineering work.
Reality
- Evidence45
- Adoption55
- Hype gap+25
- Incentives70
- Confidence45
Mid-size engineering teams can now run GitHub Advanced Security against their own repositories for 30 days without a sales call, and then they have to decide what a good result looks like. Mitch Ashley of The Futurum Group says the bottleneck was procurement.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+12
- Incentives56
- Confidence48
The Rust compiler team's first debugging survey found that 46% of respondents use a debugger, and the reasons given are about legibility rather than taste, with 74% reporting poor value rendering and 55% unable to reliably inspect a variable.
Reality
- Evidence55
- Adoption55
- Hype gap+10
- Incentives45
- Confidence58
Admins can now set filesystem, network, proxy, developer-tool and Keychain limits for Copilot sessions from inside the plugin rather than through a device-management request, and a new diagnostics tool tells them whether the policy reached the laptop.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives58
- Confidence48
Two years of compounding growth in pull request size lands on the same reviewers, while Futurum's adoption figures put AI at 40.2% in code generation against 6.2% in deployment decisions. That gap is where the queue forms.
Reality
- Evidence44
- Adoption58
- Hype gap+18
- Incentives66
- Confidence49
The PAT in dependabot.yml can finally come out, but only after an admin adds the repository to the package's Manage Actions access list, and GitHub keeps the old credential working, leaving half-done migrations indistinguishable from finished ones.
Reality
- Evidence45
- Adoption25
- Hype gap0
- Incentives55
- Confidence50
GitHub reports cost cuts of 67% on TerminalBench 2.1 and 36% on DeepSWE, all measured by GitHub, while the preview as described hands platform teams no way to set the routing policy or see which model wrote what.
Reality
- Evidence32
- Adoption15
- Hype gap+33
- Incentives74
- Confidence42
Approvals are off by default and admins can fence them to specific file paths, so the decision landing on engineering leads is how many human sign-offs a merge still needs once a machine can supply one of them.
Reality
- Evidence48
- Adoption24
- Hype gap+16
- Incentives58
- Confidence46
Earlier coverage
- Debian settles AI policy with eight-option ballot, placing review burden on developers
Product · September 1, 2026 · 1 publisher
- OpenAI starts charging some large accounts only for the jobs its model finishes
Product · August 31, 2026 · 1 publisher
- VS Code 1.135 sends an agent's work to a second model for review
Product · August 31, 2026 · 1 publisher
- GitHub will charge for Copilot seats before developers can use them
Product · August 31, 2026 · 1 publisher
- Harness gives the coding agent its own permissions and its own audit trail
Product · August 27, 2026 · 1 publisher
- CodeQL 2.26.3 treats workflow files as code, and cache poisoning as a finding you must triage
Product · August 20, 2026 · 1 publisher
- Harness hands vulnerability triage to agents, and concedes code fixes cannot keep pace
Product · August 19, 2026 · 2 publishers
- LangChain's dcode and NVIDIA's NemoClaw sell controls, not code quality
Product · August 19, 2026 · 1 publisher
- Adronite's Codistry makes token count, not context window, the axis of competition
Product · August 19, 2026 · 2 publishers
- Claude Code's 50% boost expires tonight, and your sprint capacity was a promotion
Product · August 19, 2026 · 1 publisher
- Dynatrace pays $915M for Arize, and LLM observability stops being its own category
Product · August 17, 2026 · 1 publisher