arXiv Machine Learning
techCenter
Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverabletranslating…
1 min readUnknownarXiv Digital Media
arXiv:2608.26958v1 Announce Type: new
Abstract: Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effect: larger datasets can make subtle…