Topic
Empirical laws predicting model performance from model size, data and compute, including compute-optimal training and the proposed distillation scaling law.
No current published clusters are mapped here yet.