build1 distinct publisher
Decode drags the entire model out of VRAM once per word
Prefill and decode sit on the same card and answer to different limits, which is why a GPU with more arithmetic and the same memory read rate leaves your time per output token exactly where it was.
Publishers:dev.to
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+16
- Incentives28
- Confidence