Skip to content

model

claude-haiku-4-5

Model quoted at $1.00 per 1M input and $5.00 per 1M output tokens; wins the read-heavy classification example and loses the write-heavy code example.

Known aliases

  • Anthropic Claude Haiku 4.5
  • Claude Haiku 4.5
  • claude-haiku-4-5-20251001
  • Haiku 4.5
  • MODEL_L1
  • us.anthropic.claude-haiku-4-5-20251001-v1:0

Relationships

No evidence-backed relationships are recorded.

Current stories

build1 publisher

Prompting models to carry out the task shifts false 'done' onto checks that never ran

Gemini 3.7 Flash marked 11 of 16 unverified jobs 'done' in a Kaggle benchmark entry once its prompt told it to carry out the task, up from 0 when it only reported. Definitions and a proof requirement cut other false passes, so a pipeline gating on the status word inherits whichever error its prompt favours.

Publishers:dev.to

Reality

Evidence50
Adoption
Insufficient
Hype gap+10
Incentives30
Confidence50
security1 publisher

Sophos wants SOCs to test TypeSafe's Jev for accuracy and calibration before it closes alerts

Sophos cites an independent test putting TypeSafe's Jev at 83% on Banking77 with no training examples, 10 points behind a trained classifier. Jev costs about a twelfth as much per email as Claude Haiku 4.5, so SOCs will be tempted to automate at a volume where small error rates add up.

Reality

Evidence55
Adoption
Insufficient
Hype gap+20
Incentives40
Confidence50
build4 publishers

Agent goals can spread between agents and outlive a context reset. The patch is a paragraph.

A 73-page preprint evolved instructions that jumped between coding agents and wrote themselves into the file that becomes the next system prompt. A short warning nearly stopped transmission.

Perspective Coverage

4 publishers
Builder
Builder 52%
Operator
Operator 39%
Investor
Investor 9%

Reality

Evidence68
Adoption
Insufficient
Hype gap+10
Incentives30
Confidence65
invest1 publisher

A $14.34 router matched Opus-5's score on LiteLLM's 21-task benchmark

LiteLLM's own Terminal-Bench run puts a gpt-5.4-mini classifier routing across Haiku, Sonnet and Opus at $14.34 against Opus-5's $19.74 for the same 16 solved tasks. The saving works out at 26 cents a task.

Publishers:docs.litellm.ai

Reality

Evidence45
Adoption15
Hype gap+35
Incentives80
Confidence48

Earlier coverage

  1. Ten planted bugs, about a dollar of API spend, and the case for grading the log not the answer

    Build · August 26, 2026 · 1 publisher

  2. Copilot's meter changed on June 1, and half your seats are still priced in the old unit

    Build · August 25, 2026 · 1 publisher

  3. The 97% saving was an agent failing quietly: token metrics need a completion gate

    Build · August 25, 2026 · 1 publisher

  4. AWS's phone-ordering host is really an MCP wiring diagram with no retry button

    Build · August 24, 2026 · 1 publisher

  5. Four Claude models, four surfaces, one incident: tier fallback is inside the blast radius

    Product · August 24, 2026 · 1 publisher

  6. Coding agents cost $4,125 a month because 73% of it is context you already sent

    Build · August 23, 2026 · 1 publisher

  7. Tier the models; the validation boundary is the thing you are actually buying

    Build · August 22, 2026 · 1 publisher

  8. If you can draw the flowchart before the run, you did not need the agent loop

    Build · August 22, 2026 · 1 publisher

  9. Safety fixes ship in new model versions. The regression stays with whoever pinned the old one.

    Build · August 22, 2026 · 1 publisher

  10. Physics-only world models cannot predict people, and the fix costs six pipeline stages

    Build · August 22, 2026 · 1 publisher

  11. Bedrock routing without the router Lambda: one state machine, two model calls per question

    Build · August 22, 2026 · 1 publisher

  12. Bedrock model IDs behind AppConfig flags: the swap gets cheaper, the approval gets thinner

    Build · August 22, 2026 · 1 publisher

  13. Anthropic's usage policy says no explicit content. Opus 4.6 said yes 10 times out of 10.

    Product · August 21, 2026 · 1 publisher

  14. The agent did not fail, the client did: 90 logged MCP trials and a validator that ate the calls

    Build · August 21, 2026 · 1 publisher

  15. The 21-cent model bake-off that inverted when the judge got audited

    Build · August 20, 2026 · 1 publisher

  16. A goal that writes itself into SOUL.md: agent memory is now an attack surface

    Build · August 19, 2026 · 1 publisher

  17. The cheapest model scored 10 out of 100: assistant choice is now a code-security decision

    Product · August 19, 2026 · 1 publisher

  18. A paragraph beat the agent "mind virus": reading the Anthropic-EPFL preprint as a defensive win

    Security · August 18, 2026 · 1 publisher

  19. A NIST AI RMF-mapped RAG system for $25 a month plus a third of a cent per query

    Build · August 18, 2026 · 1 publisher

  20. Claude's system prompt grew ninefold in two years. Version yours like code.

    Build · August 16, 2026 · 1 publisher

  21. Your token ratio, not the leaderboard, decides which model is cheap

    Build · August 14, 2026 · 1 publisher