NVIDIA put DOCA agent skills on GitHub and says agents using them met 100% of checklist items on 65 BlueField prompts, up from 19%. NVIDIA ran and graded that test, so the score carries over only as far as a team's own DOCA tasks resemble its prompts.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+42
- Incentives78
- Confidence55
Glow's PixelLeak report found coding agents exposed more than 13,000 private screenshots from over 300 organisations by hosting them in public GitHub repos. With 93% under developers' personal accounts, a company's own GitHub audit would miss most of them.
Perspective Coverage
3 publishers
- Builder
- Builder 42%
- Operator
- Operator 50%
- Investor
- Investor 8%
Reality
- Evidence55
- Adoption50
- Hype gap+10
- Incentives55
- Confidence60
Claude Code's new plugin eval showed one developer's skills firing in 5 of 9 relevant runs once all 89 were loaded, down from every run with one skill. The test covers three prompts in one project, but it gives teams that keep rules in skills a way to measure how often those rules get consulted.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence30
Vercel says skills.sh reached one million agent skills and nearly 280 million installs in seven months, with 375 skills taking 62% of installs. Outside the top 1.2%, skills average about 17 installs each, so the million mostly measures how cheap a skill file is to publish.
Reality
- Evidence45
- Adoption50
- Hype gap+35
- Incentives80
- Confidence55
A test published on dev.to ran one commit-message skill through Claude Code 38 times, changing only how it was described, and found that "Helps with git stuff." never fired while a file with no frontmatter always did.
Reality
- Evidence68
- Adoption
- Insufficient
- Hype gap+8
- Incentives30
- Confidence62
AgentWarden scans Markdown skills and MCP configs, exits non-zero on a rule match, and records a SHA-256 in skills.lock. It is also the project whose post documents the missing lockfile.
Reality
- Evidence38
- Adoption8
- Hype gap+10
- Incentives70
- Confidence45
An XDA Developers writer moved his desktop sticky-note prompt stash into Claude Code skill folders. Claude keeps only each skill's name and description in context and pulls in the body it judges to match.
Reality
- Evidence42
- Adoption15
- Hype gap+18
- Incentives38
- Confidence55
AWS says both coding agents it tested defaulted to Text Generation Inference and billed GPU time for each crashed deploy before pivoting to vLLM. Its answer is six editable skill files the agent reads on demand.
Reality
- Evidence44
- Adoption11
- Hype gap+21
- Incentives76
- Confidence56
The collection encodes ACMG/AMP criteria and GATK4 parameters as markdown that loads on a trigger. The 70 to 86 percent figure AWS publishes comes from head-to-head comparisons with its own unskilled agents.
Reality
- Evidence35
- Adoption15
- Hype gap+30
- Incentives80
- Confidence60
Running Claude Code's new /skill-doctor on a 40-skill setup put the standing context cost at about 3,060 tokens a turn, and the three skills it told the shop to disable were worth a fraction of the sixteen it left alone.
Reality
- Evidence58
- Adoption20
- Hype gap−10
- Incentives30
- Confidence55
Claude Code loads skills in three stages, so an idle skill charges only its description to the context window. Thomas Tartrau says the same folder runs unchanged in Cursor, and his article is the only evidence for that.
Reality
- Evidence45
- Adoption18
- Hype gap+25
- Incentives30
- Confidence55
Curated skills are files on your disk. The MCP server is a uvx proxy pinned to us-east-1. A dev.to walkthrough shows the wizard copying 23 skills into four agent trees, and the server going quiet if you press Esc.
Reality
- Evidence62
- Adoption20
- Hype gap+6
- Incentives32
- Confidence54
A preregistered pilot burned 30 agent attempts and both arms passed everything, a result that describes the three fixtures better than it describes the skill. The author reports the ceiling as his finding.
Reality
- Evidence60
- Adoption15
- Hype gap−20
- Incentives25
- Confidence55
The intel/gpu-ai-skills repo answers driver, sizing and CUDA-to-XPU questions from inside ten supported agents, which is a real cut in switching cost right up until you read the security policy and find the servers bound to 0.0.0.0.
Reality
- Evidence58
- Adoption15
- Hype gap+22
- Incentives74
- Confidence52
Six paired evaluations on a Tokyo transit server: one better route, one worse, four identical. The guidance file changed the agent's decision path without breaking anything.
Reality
- Evidence34
- Adoption11
- Hype gap+6
- Incentives52
- Confidence38
Jinguyuan's owner made his menu and wait times readable by other people's AI agents. He says it draws more press than customers, which is the interesting part.
Reality
- Evidence42
- Adoption18
- Hype gap+12
- Incentives58
- Confidence47
The new skill exists because its author found the planning stage, not the execution stage, was the thing rationing how much work he could hand to agents overnight.
Reality
- Evidence30
- Adoption16
- Hype gap+32
- Incentives68
- Confidence36
A dev.to write-up describes a statement-level dead-code tracer whose useful feature is an unresolved state. Two runs of the same skill still disagreed by 66 statements.
Reality
- Evidence32
- Adoption10
- Hype gap+15
- Incentives55
- Confidence44