Semgrep found sckit hidden in the genuine MemOS npm and PyPI packages, where it fires on Python import to scan for npm, GitHub, cloud and Slack tokens. It runs on import, not on install, so install-time scanning misses it, and anyone who imported an affected version should rotate those tokens.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence50
Tool text is fetched from the server every time an agent connects, so the description reviewed at install time can differ from the one the model reads next week. A dev.to post proposes pinning a hash of it and checking that hash in CI.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives80
- Confidence50
A dev.to checklist of ten software delivery gates for AI agents insists every check be automated, fast, blocking and observable. The gate its author calls no replacement for human review cannot meet that bar.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+40
- Incentives20
- Confidence58
DeepZero automates the hunt for exploitable Windows kernel drivers, discarding everything already on loldrivers.io before a language model rates what survives. Its maintainer reports multiple verified bugs in a driver installer corpus.
Reality
- Evidence47
- Adoption18
- Hype gap+15
- Incentives58
- Confidence52
A naive Semgrep rule fired four times across 120 generations from a 1.5B coder model and none of the four survived review, because the rule inspected try/except while the suspect default returns sat behind if guards.
Reality
- Evidence55
- Adoption15
- Hype gap−12
- Incentives20
- Confidence58
A dev.to post lists seven defect patterns in AI-generated code, and the one its author calls most distinctly AI-flavored is a dropped auth middleware that only shows up when you compare a route to its siblings.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+15
- Incentives60
- Confidence45
A study of seven agent harnesses reports 770 confirmed passes in 1,000 runs of a plugin-update attack, and no run was blocked by the model. The harness dispatches the hook, so the model has nothing to refuse.
Reality
- Evidence57
- Adoption
- Insufficient
- Hype gap+14
- Incentives56
- Confidence53
A scan of public AI repositories found the code competent and the paperwork absent, which is the harder problem. The team that ran it puts a from-zero Annex IV package at three to six months, and 2 August 2026 is behind us.
Reality
- Evidence22
- Adoption12
- Hype gap+34
- Incentives88
- Confidence30
Eleven real vulnerabilities in roughly 300 pull requests over four months. The noise was bad enough that the fix was a second agent whose only job is refuting the first.
Reality
- Evidence34
- Adoption11
- Hype gap+16
- Incentives44
- Confidence38
Codecov in 2021 and tj-actions/changed-files in March 2025 failed the same way: trusted third-party code running with pipeline privileges on a push nobody on your team made.
Reality
- Evidence52
- Adoption24
- Hype gap+18
- Incentives34
- Confidence55
A June 13 Commerce Department order cut non-US users off from two Anthropic models overnight, and OpenAI followed with limits of its own. Model access is now a sovereign risk, not a vendor term.
Reality
- Evidence22
- Adoption20
- Hype gap+45
- Incentives68
- Confidence30
Reported figures put bug rates 41% higher and security flaws 2.74x more prevalent in AI assisted code. The constraint that decides whether that matters is review throughput.
Reality
- Evidence24
- Adoption31
- Hype gap+34
- Incentives58
- Confidence21
A dev.to piece maps validate-then-commit payment flows onto a race condition class catalogued in 2006. The useful part is not the taxonomy but where the check has to live.
Reality
- Evidence56
- Adoption34
- Hype gap+24
- Incentives38
- Confidence37