ai.hackcv
论文精选 85arXiv

Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack· 好的预训练,坏的微调

Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subsequent training. We show that this assumption can fail in a full 30B mixture-of-experts training pipeline. The checkpoints that perform better after the full downstream training stack also have higher solution density, i.e., retain downstream performance under local weight perturbations.

领域:cs.AI作者:Sohir Maskey、Philipp Scholl、Jonas Knupp
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考