USENIX Security 2025 researchers found 19.7% of packages suggested by 16 LLMs were fake, and 43% of those names recurred on every re-run. Names that repeat can be registered ahead of time, so a team has to vet a suggested dependency before installing it, even when the install succeeds.
Reality
- Evidence58
- Adoption20
- Hype gap+25
- Incentives
- Insufficient
- Confidence50
Coding-agent failures cluster in the plumbing between knowing what to change and changing it. One harness builder measured what that costs and published the failure rates for Grok 4 and GLM-4.7.
Publishers:substack.com
Reality
- Evidence36
- Adoption22
- Hype gap+28
- Incentives78
- Confidence42
LectuLibre's dev.to write-up splits an EPUB on its own chapters and paragraphs, caps each chunk near 4,000 tokens, and feeds a glossary of proper nouns extracted from earlier chunks back into every translation prompt.
Reality
- Evidence58
- Adoption12
- Hype gap+22
- Incentives68
- Confidence62
A dev.to teardown ran 107 I/O-bound data engineering tasks under a sub-15-second median and a one-cent-per-task budget. What it yields is a map of where each framework's abstraction gives way once the task count climbs.
Reality
- Evidence24
- Adoption
- Insufficient
- Hype gap+38
- Incentives
- Insufficient
- Confidence27
A build report claims 87.3% root-cause accuracy across 2,400 incident scenarios and a 59% cut in mean diagnosis time. The miss rate and the denominators deserve as much attention as the headline.
Reality
- Evidence30
- Adoption15
- Hype gap+38
- Incentives62
- Confidence52
A Science Advances study finds capability tracks conformity: GPT-4 Turbo and Claude 3.5 Sonnet adopted their peers' choice between two meaningless options in groups of up to 1,000 agents.
Reality
- Evidence56
- Adoption
- Insufficient
- Hype gap+18
- Incentives32
- Confidence47