Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation
Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI) about the answer such as a reference solution, supplies dense per-token supervision to a student that never sees it. Reported gains, however, come almost exclusively from narrow, low-difficulty settings, leaving open a basic question: as a lone objective, with no reward term, does SD teach anything? We reproduce SDPO's reported gains in its easy setting, then apply the identical setup to difficult tasks and find that it does not. Across question answering, mathematics, coding, and multi-turn agentic tool use, across reasoning modes, model sizes, and forms of PI, and under both the SDPO and OPSD recipes, the per-t
质疑自我蒸馏在复杂任务中的有效性
- 核心方法
- 通过对不同任务类型、模型大小和特权信息形式进行实验,研究自我蒸馏的效果
- 适合谁读
- 研究者
- 要解决的问题
- 研究自我蒸馏是否在没有奖励项的情况下,仅依赖特权信息就能有效教学
- 关键实验
- 实验涵盖问答、数学、编程和多轮代理工具使用等任务,使用不同模型大小和特权信息形式
- 主要贡献
- 揭示了自我蒸馏在简单任务中的成功不易推广到复杂任务
- 意义与局限
- 对自我蒸馏方法的有效性提出了质疑,影响未来研究方向的选择;局限在于未能详细说明失败原因