product1 distinct publisher
The 150,000-to-1 gap: scaling is a workaround for a learning mechanism nobody has found
Llama 3.1 was pretrained on 15 trillion tokens. A preteen manages on about 100 million words. Nobody knows how the child does it, and easily available web data may run out in the 2030s.
Publishers:technologyreview.com
Reality
- Evidence54
- Adoption20
- Hype gap+12
- Incentives32
- Confidence52