build1 publisher
Claude Code's plugin eval spends six agent runs per case to measure a plugin's lift
Anthropic's plugin eval command runs every case with the plugin loaded and again without it and prints the delta, though the CI threshold it documents still gates on each case's absolute score.
Publishers:dev.to
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence58