ai.hackcv
论文精选 65arXiv

Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection· 重试、切换或弃权?学习策略感知的工具使用政策

Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail transiently, persistently, or silently. Robust recovery therefore requires more than repeated retries: an agent may need to retry the same path, switch to an alternative, or recognize that no viable path remains. We present BENCH2ROBUST, a framework that converts failure-free tool-use benchmarks into controlled stochastic environments with scenario-controlled solvability, where episodes explicitly require retrying, switching, or stopping after available paths are exhausted. We use BENCH2ROBUST to study two complementary interventions: structured runtime recovery context through Bayesian Tool Memory (BTM), and curriculum-controlled reinforcement learning. A

AI 解读论文

解决工具使用代理面对工具失败时的恢复策略。

核心方法
通过BENCH2ROBUST框架将无故障的工具使用基准转换为可控随机环境,并结合贝叶斯工具记忆(BTM)和课程控制的强化学习来研究恢复策略。
适合谁读
研究者、工程师
要解决的问题
现有的工具使用代理在工具失败情况下缺乏有效的恢复策略,影响其鲁棒性。
关键实验
在多个工具使用任务上进行了实验,包括不同类型的工具失败,验证了BTM和课程控制的强化学习的有效性。
主要贡献
提出了一种新的框架BENCH2ROBUST,可以更好地评估和提高工具使用代理在工具失败情况下的鲁棒性和恢复能力。
意义与局限
为提高工具使用代理的鲁棒性提供了新的方法论;但在真实世界应用中的效果仍需进一步验证。
领域:cs.AI作者:Chaoran Chen、Vy Nguyen、Ziji Zhang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考