ai.hackcv
论文精选 65arXiv

TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards· TRACE: 通过合成奖励训练代理进行因果探索

Reinforcement learning with verifiable rewards (RLVR) has advanced language-model reasoning in domains such as mathematics and code, where objective answers are inexpensive to check. Diagnostic reasoning over complex data lacks this advantage: establishing the true cause of an anomaly often requires costly expert investigation and may remain ambiguous after the fact. We ask whether this asymmetry of verification can instead be engineered. We sample an intervention, inject it into a controlled simulator, and generate the observations it would produce. The hidden intervention provides an oracle label and objective reward, while the agent must still investigate noisy, confounded, and distributed evidence. We instantiate this approach in TRACE, a digital-advertising diagnostic environment with

AI 解读论文

通过合成奖励训练代理进行因果探索,解决复杂数据诊断中的验证难题。

核心方法
通过在受控模拟器中注入隐藏干预并生成相应的观察结果,创建合成奖励,使代理在处理有噪声、混杂和分布的证据时能够学习因果关系。
适合谁读
研究者 / 工程师
要解决的问题
在复杂数据诊断中,确定异常的真正原因通常需要专家调查且成本高昂,而这种方法难以验证。
关键实验
在 TRACE 环境中进行了多项实验,验证了合成奖励对代理学习因果关系的有效性。
主要贡献
提出了 TRACE 环境,用于数字广告诊断,能够通过合成奖励高效训练代理进行因果探索。
意义与局限
这种方法为提升复杂数据诊断的自动化水平提供了新的思路,但在真实世界应用中仍需进一步验证。
领域:cs.AI作者:Rui Sun、Zhan Shi、Bing He
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考