arXiv CS AI
techCenter
When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillationtranslating…
arXiv:2608.27960v1 Announce Type: new
Abstract: On-policy distillation (OPD) has recently emerged as a popular post-training paradigm for large language models (LLMs), providing an efficient way to transfer the knowledge and capabilities of teacher models into student models…
Keywords#On-Policy Distillation#Teacher#models#Guidance#Guidance Misleads