ai.hackcv
论文精选 65arXiv

Rethinking On-Policy Distillation of Large Language Models II: One Training Example· 重新思考大语言模型的单样本策略蒸馏

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure \emph{state coverage}, the fraction of the states full-data OPD visits that a query set's rollouts reach. A single query already reaches \(71.5\%\), most of it within the first 100 steps. Adding semantically distinct queries raises cov

AI 解读论文

单样本策略蒸馏在大语言模型训练中表现出显著效果。

核心方法
通过分析训练过程中的状态覆盖度和学生模型与教师模型的对齐速度,研究单样本策略蒸馏(One-shot OPD)。
适合谁读
研究者
要解决的问题
理解单个训练样本在大语言模型的策略蒸馏过程中的作用。
关键实验
实验展示了单个样本和语义上不同样本在不同任务域和模型家族中的效果;特别地,单个样本在前100步内即达到71.5%的状态覆盖率。
主要贡献
证明了单个样本在OPD过程中可以持续改进学生模型,并恢复大多数全数据OPD的性能增益。
意义与局限
为大语言模型的高效训练提供了新的视角;不过,研究主要集中在特定条件下,其普遍适用性还需进一步验证。
领域:cs.AI作者:Zixuan Fu、Bingxiang He、Yuxin Zuo
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考