Skip to content

model

deepseek-v4-flash

Model reported as deterministic (3 of 3 identical responses) with code_first unknown rates of 0.94 in English and 0.54 in Russian prompts.

Known aliases

  • deepseek/deepseek-v4-flash
  • deepseek_v4
  • DeepSeek-V4
  • DeepSeek V4 Flash
  • DeepSeek V4-Flash
  • deepseek-v4-flash
  • DeepSeek-V4-Flash-0731
  • DeepSeek-V4-Flash-Max
  • V4 Flash
  • V4-Flash

Relationships

No evidence-backed relationships are recorded.

Current stories

build4 publishers

DeepSeek open-sources the harness, then raises the price of the model

Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.

Perspective Coverage

4 publishers
Builder
Builder 51%
Operator
Operator 31%
Investor
Investor 18%

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives70
Confidence58
build3 publishers

DeepSeek's new encoder-decoder splits inference into an 8B prefill and a 16B decode

V4.1-Flash retires the V4 Pro line and carries two active-parameter counts, 763B total with 8B on input tokens and 16B on output, so one sizing number no longer covers both phases of a request. Baseten had it running on day zero.

Publishers:businesstimes.com.sgdev.tolatent.space

Perspective Coverage

3 publishers
Builder
Builder 40%
Operator
Operator 28%
Investor
Investor 32%

Reality

Evidence60
Adoption35
Hype gap+25
Incentives40
Confidence58
build5 publishers

Alibaba ships the Qwen4 architecture as open weights before the flagship exists

Qwen3.8-Flash-Next puts 36 Gated DeltaNet layers and 12 sparse-attention layers on Hugging Face, which means the retrieval budget Qwen4 will inherit is something you can measure against your own traces now.

Perspective Coverage

5 publishers
Builder
Builder 52%
Operator
Operator 28%
Investor
Investor 20%

Reality

Evidence58
Adoption52
Hype gap+32
Incentives76
Confidence71

Earlier coverage

  1. DeepSeek V4 moves the coding-model decision into the finance column

    Build · September 1, 2026 · 1 publisher

  2. Thirty-nine retries fit inside the price gap between GLM-5.3-Flash and Opus 4.8

    Build · August 31, 2026 · 1 publisher

  3. Peak-hour pricing pushes DeepSeek's new vision model past Gemini on the invoice test

    Build · August 31, 2026 · 2 publishers

  4. Five buyers now owe Nvidia $44bn of the demand its guidance promised

    Invest · August 27, 2026 · 1 publisher

  5. Harness choice moved token use 83-fold with the model held constant

    Build · August 27, 2026 · 1 publisher

  6. Stealth is now a launch strategy: Zhipu's Ox Alpha topped the charts before it had a name

    Product · August 26, 2026 · 1 publisher

  7. Ox Alpha passes the Xinjiang test and fails on Xi: seven topics, 83 points apart

    Build · August 25, 2026 · 1 publisher

  8. DeepSeek V4 doubled its OpenRouter token share, and the bill it displaced was ~130x bigger

    Invest · August 25, 2026 · 1 publisher

  9. A 284B model at 25 tokens a second on one 5090, and 192 GiB of DDR5 doing the quiet part

    Build · August 22, 2026 · 1 publisher

  10. A 284B model in 3.2GB of RAM turns sparse MoE into a disk bandwidth problem

    Build · August 21, 2026 · 1 publisher

  11. Cost per shipped feature, not the leaderboard: one CTO cut a $14k model bill by $9k

    Build · August 19, 2026 · 1 publisher

  12. Re-baseline AI procurement on cost per completed task, not dollars per million tokens

    Leadership · August 18, 2026 · 1 publisher

  13. Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill

    Build · August 18, 2026 · 1 publisher

  14. A Beijing bar gives away DeepSeek tokens. Your per-token price list is the collateral damage.

    Product · August 17, 2026 · 1 publisher

  15. Four months of A100 bills say self-hosting is a utilization bet, not a cost saving

    Build · August 16, 2026 · 1 publisher

  16. DeepSeek V4 Flash costs a tenth as much and passes 53.8% of agent tasks

    Invest · August 16, 2026 · 1 publisher

  17. The zero-false-accept memory result belonged to the proxy, not the model

    Build · August 14, 2026 · 1 publisher