science1 publisher
Why LLM output is hard to reproduce: it's not just concurrency and floating point
Thinking Machines Lab argues that temperature-zero inference varies run to run not because GPUs are chaotic but because kernels change their reduction order with load, which makes bit-identical output an engineering target.
Publishers:thinkingmachines.ai
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+12
- Incentives60
- Confidence55