One developer moved about a dozen agent rules out of an instructions file and into Claude Code PreToolUse hooks after watching them slip in long contexts. Each hook blocks a tool call before it executes, though it guards only the command shapes its pattern matches.
Reality
- Evidence35
- Adoption8
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
Coding-agent teams should move their top three irreversible actions behind checks outside the model, a dev.to guide argues, because prompt rules are advice. The argument rests on how prompts are wired, and the post's own release advice weakens one of its three objections.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
The startup says Anchor 3.0 caught more than 90 percent of violations in a test it published itself, at under a 500th of the cost of a frontier call, and the messages it misses stay the firm's problem.
Reality
- Evidence30
- Adoption16
- Hype gap+32
- Incentives82
- Confidence58
An engineer measured TypeSafe AI's Jev decision model on 10 public datasets and published the raw results. The two largest movements came from decomposing and staging the questions, with the model build held constant.
Reality
- Evidence62
- Adoption12
- Hype gap−5
- Incentives
- Insufficient
- Confidence55
A 29-session test on Claude Code v2.1.273 ran the same protected-directory rule two ways, as prose in CLAUDE.md and as a PreToolUse hook. Both held on a plain task. The comparison that separates them rests on four runs.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence55
Part 4 of a dev.to enterprise AI series puts a person in front of the write and send connectors and leaves read paths alone, with the refusal enforced by the tool layer instead of a sentence in the prompt.
Reality
- Evidence38
- Adoption15
- Hype gap−8
- Incentives25
- Confidence50
Part 14 of a dev.to series describes what its author calls "a support agent that can work out whether you're owed a refund, and cannot give you one", with the eligibility rules pinned to the JDK by a test.
Reality
- Evidence50
- Adoption4
- Hype gap−10
- Incentives28
- Confidence45
On the platform behind taabi Nexus, the model's only output is a validated rule document, and the code that dials a driver's intercom sits behind a seven-day replay, a role check and one person's approval.
Reality
- Evidence58
- Adoption24
- Hype gap−12
- Incentives62
- Confidence55
The MIT-licensed framework gives agents one command grammar for Salesforce, ServiceNow, DocuSign and Agentforce. The safety metadata a host checks before running a command comes from provider authors the project leaves unnamed.
Reality
- Evidence46
- Adoption11
- Hype gap+14
- Incentives66
- Confidence44
Newer models ship with advice to keep the guidance file under 200 lines. The analyst who wrote those 1,042 lines argues each one logs context the model lacked, and that the real debt is a mechanical rule left sitting in prose.
Reality
- Evidence42
- Adoption22
- Hype gap−12
- Incentives38
- Confidence55
A developer instrumented 89 coding sessions to test whether rules the agent had already read changed what it did. Enforcement only started working once the rule was rewritten into something a hook could see.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap−10
- Incentives30
- Confidence55
Writing on dev.to, Ben McCarthy describes rebuilding the guest-messaging agents across his 23 holiday lets around a typed per-request view of state, with world facts held apart from the policy about what the agent may promise.
Reality
- Evidence32
- Adoption20
- Hype gap+15
- Incentives35
- Confidence55
Moving a rule out of CLAUDE.md and into an exit code stopped the reoffending, but the 46-line guard that replaced it decides handoff syntax by regex, without knowing whether the command it inspects ever returns.
Reality
- Evidence45
- Adoption8
- Hype gap+38
- Incentives60
- Confidence52
The 90 percent token saving is one engineer's own usage on his own repos, but the enforcement mechanism underneath it copies cleanly: a line-count threshold that returns a block, with a pipe-shaped hole in it.
Publishers:engineering.atspotify.com
Reality
- Evidence44
- Adoption18
- Hype gap+30
- Incentives72
- Confidence52
An adversarial self-audit check that blocks with exit 2 is a useful nag in a terminal. Under launchd it becomes a failed job. The fix here keys off session provenance rather than softening the rule.
Reality
- Evidence52
- Adoption10
- Hype gap+12
- Incentives28
- Confidence55
The hooks documentation says a matcher on an event without matcher support is ignored rather than rejected. A Bash-scoped guard on UserPromptSubmit therefore fires on every prompt, and nothing at startup or in --debug says so.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap−10
- Incentives55
- Confidence58
The prompt only ever guarded sessions someone was watching, so the rules moved into a hook that runs in bypass and headless mode, refuses a force push by naming the sanctioned path, and self-tests on 59 synthetic calls.
Reality
- Evidence48
- Adoption14
- Hype gap+12
- Incentives22
- Confidence55
Allow and deny globs in settings.json cannot separate git status from git push --force, so the real enforcement moves into a PreToolUse hook that reads the command text, exits 2, and tells the model why it was refused.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+28
- Incentives62
- Confidence55
Command substitution inside the echo reset the exit status to zero, so a Stop hook reported success for a full day while DM replies sat at zero on two platforms. The test that catches this takes a minute to write.
Reality
- Evidence66
- Adoption18
- Hype gap+12
- Incentives55
- Confidence61
Agent Package Resolution, now in preview, binds Claude Code and Cursor to Artifactory. Two of its three enforcement layers still run where the agent does.
Reality
- Evidence34
- Adoption14
- Hype gap+42
- Incentives86
- Confidence44