build1 publisher
Folding 92 layers into one Pallas call put Kimi K3 at 709 tokens a second on TPU v7
Inferact measured 709 output tokens a second on 16 Ironwood chips against 452 on 16 GB200s, at low concurrency with speculative decoding. The engineering worth reading is the hand-written memory schedule underneath.
Publishers:runtimewire.com
Reality
- Evidence45
- Adoption22
- Hype gap+20
- Incentives78
- Confidence58