leadership1 publisher
DeepSeek's V4.1-Flash reads a million-token prompt on 8B active parameters
The model card for this 552B-parameter Mixture-of-Experts release puts the global KV cache at 890 bytes per token, about a quarter of the previous Flash generation, and every figure in it is DeepSeek's own.
Publishers:huggingface.co
Reality
- Evidence58
- Adoption20
- Hype gap+18
- Incentives82
- Confidence60