ai.hackcv
论文精选 85arXiv

Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference· 不要放弃Dropout:优化层稀疏性以提高LLM训练和推理效率

Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout - particularly layer dropout - has largely disappeared from large language models (LLMs) pre-training recipes. While some prior work has reported that dropout can degrade accuracy, no comprehensive study has quantified, let alone mitigated, this effect. In this study, we show that layer dropout should be used in state-of-the-art LLM training, establishing best practices and scaling analysis for both training and post-training benefits. Concretely, with optimal layer distribution, time schedule, and optimizer hyperparameters, we observe that at the same train

领域:cs.AI作者:Mostafa Elhoushi、Alex Pretko、Nolan Dey
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考