Skip to content

Topic

LLM Inference Cost

Expense of running LLM API calls, driven by per-token pricing, context length, retries, and caching; tracked per request or per million tokens.

Current stories

build3 publishers

DeepSeek's new encoder-decoder splits inference into an 8B prefill and a 16B decode

V4.1-Flash retires the V4 Pro line and carries two active-parameter counts, 763B total with 8B on input tokens and 16B on output, so one sizing number no longer covers both phases of a request. Baseten had it running on day zero.

Publishers:businesstimes.com.sgdev.tolatent.space

Perspective Coverage

3 publishers
Builder
Builder 40%
Operator
Operator 28%
Investor
Investor 32%

Reality

Evidence60
Adoption35
Hype gap+25
Incentives40
Confidence58
build1 publisher

LangChain drops about 4,000 base input tokens from every default Deep Agents turn

The harness lost its hidden system prompt, 43% of its builtin tool descriptions and its todo list middleware. LangChain's own footnote says reward confidence intervals span zero for every model tested, so the evals settle the token saving more firmly than the quality.

Publishers:langchain.com

Reality

Evidence58
Adoption30
Hype gap+18
Incentives82
Confidence46

Earlier coverage

  1. A timestamp in the system prompt turns prompt caching into a 25% surcharge

    Build · September 12, 2026 · 1 publisher

  2. A demo-to-production LLM checklist prices 50,000 daily calls at two different rates

    Build · September 11, 2026 · 1 publisher

  3. AgentJIT compiles a traced agent run into deterministic Python after one warmup call

    Build · September 11, 2026 · 1 publisher

  4. A 1.5x per-token price still bought a 25 percent cheaper correct answer in AWS's benchmark

    Build · September 11, 2026 · 1 publisher

  5. FrugalGPT fits a fresh triage rule for every dataset and task it is tested on

    Build · September 10, 2026 · 1 publisher

  6. Mistral Small 3.2 cuts the same text into 547 Polish tokens and 377 English ones

    Build · September 10, 2026 · 1 publisher

  7. Re-running five Terminal-Bench-Science tasks at $12 each leaves Fable 5.1 passing one

    Build · September 10, 2026 · 1 publisher

  8. A forked Claude Code skill ran cat CLAUDE.md to reach the secret its prompt withheld

    Build · September 9, 2026 · 1 publisher

  9. A 33-run sweep prices OpenAI's reasoning_effort ladder at 2.3x for identical answers

    Build · September 7, 2026 · 1 publisher

  10. Compaction that cut tool output 38.4% pushed the bill up 6.8%

    Build · September 2, 2026 · 1 publisher

  11. Five meters, one minute: why voice agent budgets should be priced per outcome

    Product · August 18, 2026 · 1 publisher

  12. Your token ratio, not the leaderboard, decides which model is cheap

    Build · August 14, 2026 · 1 publisher