ai.hackcv
论文精选 60arXiv

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning· 解构离策略比率:异步强化学习中的熵缩放信任区域

Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause policy collapse. Existing methods typically retain or discard tokens based solely on the magnitude of their importance ratios, applying the same threshold uniformly across token positions. In this work, we reveal that the natural scale of the importance ratio varies systematically with token entropy. Under asynchronous dynamics, this entropy-ratio scaling dictates two distinct phenomena: at low entropy, the inherent train-inference discrepancy is drastically amplified into substantial sampling noise; at high entropy, in-flight weight updates natural

AI 解读论文

解释离策略比率与熵的关系,提出异步强化学习的信任区域方法

核心方法
通过分析重要性比率与token熵的关系,提出熵缩放信任区域方法,动态调整采样阈值
适合谁读
研究者
要解决的问题
异步强化学习中,由于数据滞后导致的优化不稳定和策略崩溃问题
关键实验
未提供
主要贡献
揭示了离策略比率的自然尺度与token熵的系统性变化,提出了改进异步优化稳定性的方法
意义与局限
该方法有助于提高异步强化学习的稳定性,减少策略崩溃的风险,但可能需要针对不同任务进一步验证
领域:cs.AI作者:Guanqun Zhao、Zijun Xie、Binbin Zheng
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考