One 8,192-token session on a 27B model holds 512 MB of key-value cache, or 64 KB for every token generated. How many of those sessions fit in free VRAM sets serving concurrency, and paging decides the waste.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+24
- Incentives45
- Confidence48
A three-part source walkthrough traces vLLM V1 from generate() to the CUDA boundary. The useful finding is structural: the scheduler that spends your token and KV-block budgets runs in a different process from the one you are timing.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap−5
- Incentives30
- Confidence60
A ByteByteGo teardown of the three open-weight serving engines describes three different request disciplines. Queuing and cache reuse decide what a GPU can serve before the model matters.
Reality
- Evidence26
- Adoption
- Insufficient
- Hype gap+32
- Incentives44
- Confidence34
A CUDA and ROCm tutorial ends on the unglamorous half of Mixture of Experts: two all-to-all collectives per layer, optimizer state pushed onto PCIe, and launch overhead you have to design away.
Reality
- Evidence24
- Adoption
- Insufficient
- Hype gap+42
- Incentives32
- Confidence58
A dev.to explainer on local inference benchmarks makes a point worth pinning up: tokens per second is a function of how many users you tested with, not a property of the hardware.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+14
- Incentives42
- Confidence41