build1 publisher
DeepSeek's new encoder-decoder splits inference into an 8B prefill and a 16B decode
V4.1-Flash retires the V4 Pro line and carries two active-parameter counts, 763B total with 8B on input tokens and 16B on output, so one sizing number no longer covers both phases of a request. Baseten had it running on day zero.
Publishers:latent.space
Reality
- Evidence58
- Adoption52
- Hype gap+25
- Incentives65
- Confidence55