ai.hackcv
论文精选 60arXiv

Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning· 学习使用工具:强化学习以实现工具集成的数学推理

Current large language models (LLMs) increasingly benefit from external tool integration, especially for tasks requiring reliable computation and verification. Motivated by this, we study calculator tool calling for improving mathematical reasoning on the Countdown task. We first analyze reasoning failures and find that calculation errors account for a substantial portion of incorrect responses. We then construct supervised fine-tuning datasets to teach the model useful tool-use patterns and how to interpret returned outputs. Building on this tool-formatted policy, we apply several on-policy reinforcement learning methods, including RLOO, RLOO++, GRPO, and DAPO, using automatically verifiable final-answer rewards. To enable a more reliable evaluation, we construct a fresh 1,024-problem hel

AI 解读论文

强化学习结合工具使用,提升数学推理能力

核心方法
通过构建监督微调数据集,教授模型如何有效使用计算器工具及解释结果,并应用多种强化学习方法
适合谁读
研究者 / 工程师
要解决的问题
解决现有大型语言模型在执行数学任务时的计算错误问题
关键实验
构建了1,024个新问题的数据集用于评估模型性能
主要贡献
提出并验证了使用强化学习方法结合外部计算工具,改进数学推理准确性的技术路径
意义与局限
提高了数学推理任务中模型的可靠性与准确性,推动了大型语言模型与外部工具集成的研究
领域:cs.AI作者:Minghui Xu、Zi Wang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考