build1 publisher
llama.cpp takes roughly half an hour to reach first token on an RTX 5090
Alibaba's 27B model fits a 32GB card with 15GB to spare. Tom's Hardware still had to pick between llama.cpp's full 262K window at half-hour prefill and a supported vLLM deployment capped at 32K.
Publishers:tomshardware.com
Reality
- Evidence60
- Adoption38
- Hype gap+15
- Incentives55