Addy Osmani's Opus 5.5 guide for Anthropic says to delete 'think carefully' lines and give each task a finish line and one stop condition. Its sturdier advice covers long Claude Code runs, where a CLAUDE.md rule tells the model when to keep going and when to stop.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence35
One developer running Claude Code on 80-plus microservices enforces a never-commit rule, once pasted into 18 prompt files, with a permission deny rule. An audit found the agent's settings file allowing the git commands those prompts banned in capitals.
Reality
- Evidence35
- Adoption15
- Hype gap+5
- Incentives
- Insufficient
- Confidence40
Plain Claude Code, with no MCP server or skill, matched AWS's and draw.io's official diagram tools in a five-setup test published on dev.to. Every setup improved as the author kept adding instructions, and the finding covers one architecture on Claude Opus 5, graded by the author.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
Gemini 3.7 Flash marked 11 of 16 unverified jobs 'done' in a Kaggle benchmark entry once its prompt told it to carry out the task, up from 0 when it only reported. Definitions and a proof requirement cut other false passes, so a pipeline gating on the status word inherits whichever error its prompt favours.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence50
Misspelling 70% of a prompt's words left Claude models' scores unchanged across about 4,900 test sessions, but one wrong punctuation mark cost 8 to 23 points. Both breaks that stuck erased the line between instruction and data, so delimiters deserve the review time that spelling gets.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
Developer hram had a coding agent build a Kotlin messenger for Android, iOS, desktop and web from a 33 KB master prompt split into five step prompts. Every prompt is public and unedited, so teams can test the plan-first method against the code it produced.
Reality
- Evidence50
- Adoption5
- Hype gap0
- Incentives
- Insufficient
- Confidence55
The model documentation concedes Astra asks for clarification more often than GPT-5.6 Sol and sometimes stops early. Every remedy on the page is more instruction text, loaded into the same context the docs blame for stalls.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+10
- Incentives50
- Confidence65
A developer's free resume builder bans its AI from adding nine named kinds of fact and rejects any rewrite that changes the bullet count. The live check covers only the count, so a claim invented in words inside a single bullet still reaches the user, who is told to read every result.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives35
- Confidence45
The measured half is a word count over 24 runs on three everyday questions. One eval arm's result shifted between two patch releases of Claude Code, and the transcript layer rides on hooks the CLI does not document.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+12
- Incentives70
- Confidence55
In a 30-call test suite for a home services intake agent, guardrail G4 told the model to stop collecting fields the moment it heard a gas smell and never told it when to resume. The dispatcher got the result.
Reality
- Evidence60
- Adoption8
- Hype gap−10
- Incentives35
- Confidence50
Clusterflick sorts 450 to 600 unmatched London cinema listings a day into ten categories. Its maintainer spent an evening rebuilding that call around typed questions and exclusion lists, with the feature-length comparison done in his own code.
Reality
- Evidence55
- Adoption15
- Hype gap+5
- Incentives35
- Confidence45
A dev.to post puts the fragile layer of a multi-tool agent in the choice between two tools that both fit the request, and argues that better tool descriptions plateau because the ambiguity belongs to the request itself.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+12
- Incentives22
- Confidence34
In a dev.to walkthrough, miruky uses one bad rstrip call to separate five layers of agent engineering. Which layer do you change when a repair passes all three examples and still accepts 1e3s?
Reality
- Evidence70
- Adoption
- Insufficient
- Hype gap−5
- Incentives20
- Confidence58
In a paper posted on 13 September, nine models across the Claude, GPT and Gemini families edited already-optimal EffiBench solutions in every trial. One added sentence in the prompt recovered 20 refusals.
Reality
- Evidence45
- Adoption18
- Hype gap+15
- Incentives35
- Confidence55
Ivan Burazin told Business Insider that Daytona's best agent operators are former people managers, and the org he described runs about 80 agents behind 16 engineers who no longer write code themselves.
Reality
- Evidence42
- Adoption20
- Hype gap+30
- Incentives80
- Confidence58
Toolmetry edited only the strings that tell an agent what each MCP tool does and how to call it. On a single published run, its three test servers finished between 96.4 and 100 percent for $4 of API spend.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+55
- Incentives55
- Confidence38
A dev.to breakdown prices a fine-tune at £20,000 to £80,000 for a mid-sized UK build, most of it domain-expert labour on the dataset, and argues that the request usually describes a gap retrieval closes.
Reality
- Evidence36
- Adoption
- Insufficient
- Hype gap+18
- Incentives58
- Confidence42
A retired production prompt measured 56,000 tokens on every call, ninety percent of it accumulated bans. Its author says the rule-per-failure loop looked like progress and produced committee prose.
Reality
- Evidence38
- Adoption12
- Hype gap+20
- Incentives40
- Confidence50
Thibault Schrepel randomised his Vrije Universiteit Amsterdam law students into ban, unguided ChatGPT and trained arms, ran the design in 2024 and again in 2025, and the trained arm's lead was gone by the second year.
Reality
- Evidence35
- Adoption20
- Hype gap+30
- Incentives40
- Confidence55
A dev.to post argues that the rules worth committing to .github/copilot-instructions.md are the ones a linter cannot check and the model cannot infer, such as money being Decimal and datetime.utcnow() returning a naive datetime.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+12
- Incentives45
- Confidence55
Earlier coverage
- Thirty SKILL.md files cost about 3,000 tokens until one of them activates
Build · September 11, 2026 · 1 publisher
- FrugalGPT fits a fresh triage rule for every dataset and task it is tested on
Build · September 10, 2026 · 1 publisher
- Only one of three model backends can refuse a wrongly shaped JSON response
Build · September 8, 2026 · 1 publisher
- ZeroShot Studio moved its rewrite veto out of the prompt and into a pre-write code hook
Build · September 8, 2026 · 1 publisher
- Deleting an example beat banning it across three rebuilds of a 680-line prompt
Build · September 7, 2026 · 1 publisher
- Requiring a reproduction path stops a review agent from quadrupling the feature
Build · September 6, 2026 · 1 publisher
- An unsupervised agent loop billed $38 before anything in the system said stop
Build · August 30, 2026 · 1 publisher
- Anthropic's /eli5 skill: two lines of instruction, one HTML file, no room for precision
Build · August 24, 2026 · 1 publisher
- Before you buy another GPU, check num_ctx and the rope base
Build · August 22, 2026 · 1 publisher
- Ottawa funds AI literacy; the part employers actually need is the audit habit
Science · August 20, 2026 · 1 publisher
- Repetitive LLM output is three separate defects, and most teams fix one and stop
Build · August 18, 2026 · 1 publisher
- The payload is rebuilt every turn, so stop treating your prompt as a shipped artifact
Build · August 15, 2026 · 1 publisher
- Agent reliability is a harness problem, not a prompt problem
Build · August 15, 2026 · 1 publisher