Loading today’s stories
other
Preprint proposing a distillation scaling law and compute-optimal recipes for existing versus newly trained teachers, based on a controlled study with 143M to 12.6B parameter models.
No evidence-backed relationships are recorded.
No current published clusters are mapped here yet.