Build1 publisher3 min readPublished
Anthropic's Opus 5.5 guide swaps 'think carefully' for finish lines and stop rules
Addy Osmani's Opus 5.5 guide for Anthropic says to delete 'think carefully' lines and give each task a finish line and one stop condition. Its sturdier advice covers long Claude Code runs, where a CLAUDE.md rule tells the model when to keep going and when to stop.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Osmani published the guide on the Anthropic engineering blog last week, and it reached the Hacker News front page.
- Anthropic tested dropping a 'think carefully' line in a chat product of its own: replies began faster, and quality showed no clear decline.
- Early testers ran Opus 5.5 on hours-long coding tasks with little oversight, and its largest gains over Opus 5 came in pushing changes through big repositories until tests pass.
- The guide warns that Opus 5.5 sometimes halts mid-task with a summary naming the next step, or an offer to continue, without taking that step.
- It recommends keeping the task list in TASKS.md, because Claude Code summarizes older turns once long runs fill the context window.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Teams with reasoning cues in saved instructions pay any added delay on every message they send, so editing those files is the cheapest change the guide asks for.
- decision Writing a prompt now means choosing in advance the one failure worth an interruption, with everything else downgraded to a status note.
- exposure Under a keep-going rule a person sees fewer pauses in which to catch a bad step, so deletes and force-pushes depend on the carve-out and the tool's permission prompts.
Opus 5.5 thinks before every reply and picks how much thinking the task deserves, according to a dev.to write-up of the guide [4]. A "think carefully" line asks the model for something it already does. The write-up does not give a latency figure for Anthropic's chat-product test or say how quality was judged [3].
Treat that test as a claim about a chat workload. For it to transfer, a team's tasks would need to resemble chat traffic. The quality check would also need to include the hardest prompts, where a nudge toward more reasoning was most likely to matter. The write-up's author says multi-week experiments in live repos confirmed the guide's core claim, and adds that the numbers and recommendations are the guide's [18]. I'd still compare outputs with and without the line on a team's own saved prompts before deleting it everywhere.
The prompt shape is the better engineering in the guide. Osmani's example puts the whole task in one message: "Migrate the payment endpoints from the old client to the new one." [6] The finish line follows: "Done means: every endpoint uses the new client, the old client is deleted, and the test suite passes." [7] Then one stop condition: "Stop and ask me only if a test fails for a reason you can't explain." [8] Each clause of that finish line is something the model can check on its own: a search for the old client, a deleted file, a passing test run. The shape fits the long, lightly supervised repository work where the guide places the model's biggest gains [9].
For the mid-task stall, the write-up points to a two-line CLAUDE.md rule from Osmani [10]. The first line reads: "When a step doesn't need my input, keep going. Put status notes in the same message as your next action." [11] The second sets the exits: "Stop and ask only when you can't continue without me, or before anything destructive: deleting data, force-pushing, or changing anything outside this repository." [12] For a single stall, the fix is the lowest-tech one in the guide: reply "continue" [13].
A keep-going rule means fewer stops. So the write-up keeps a human checkpoint before anything risky or hard to undo, and leaves permission prompts on for destructive commands [14]. I think that split is right. The CLAUDE.md line is an instruction the model interprets. The permission prompt is a setting in the tool, and it still applies when the model misjudges what counts as destructive.
For audits and migrations across a large codebase, Osmani suggests a fan-out: "Give each service to its own subagent. When a subagent reports back, check its evidence before you accept it." [15] The run ends in one table listing each service, whether it is affected, and the evidence [16].
What to watch
- Whether Anthropic publishes the latency and quality figures behind its chat-product test, or repeats the test on agent sessions in Claude Code.
- Whether users still see runs end with 'Want me to continue?' after adding Osmani's keep-going rule to CLAUDE.md.
- Reports of Opus 5.5 deleting data, force-pushing, or editing outside the repository while running under a keep-going rule.