build1 publisher
Chip Huyen puts inference at 10 to 100 times a model's training compute
Her P99 keynote sets hardware aside because most teams cannot change it. That leaves the weights and the serving layer, and her own goodput example shows where the seconds in a response actually go.
Publishers:thenewstack.io
Reality
- Evidence55
- Adoption25
- Hype gap+15
- Incentives55
- Confidence55