ai.hackcv
论文精选 65arXiv

On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification· 自改进代理的脆弱性

Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods have been critically overlooked. In this work, we conduct a comprehensive re-evaluation of two memory-based methods, broadening the scope of evaluation along two axes: (1) including multiple runs to quantify variance, and (2) randomly shuffling the tasks to investigate the effect of task order. Through these experiments, we make two observations that expose the fragility of current methods: First, agent evaluation is inherently noisy in complex environments and on multi-step tasks, and stacking a self-improving loop on top can further amplify this noise

AI 解读论文

评估自改进代理的不稳定性及其对任务顺序的敏感性

核心方法
通过多次运行和随机任务顺序重新评估记忆基自改进代理,分析其性能变化和脆弱性
适合谁读
研究者和工程师
要解决的问题
解决自改进代理在复杂环境和多步骤任务中的可靠性问题
关键实验
包括多次运行以量化方差,以及随机任务顺序的影响实验
主要贡献
揭示了自改进代理在评估中的噪声问题和对任务顺序的敏感性
意义与局限
提高了对自改进代理可靠性的认识,提醒研究者在设计和评估时需考虑环境复杂性和任务顺序的影响,但也限于特定评估方法
领域:cs.AI作者:Qinyuan Ye、Yu Li、Yada Pruksachatkun
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考