ai.hackcv
论文精选 65arXiv

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development· 超越最终得分:对长期AI研究开发代理的系统评估

Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers:

AI 解读论文

超越分数评估长期AI研究代理的行为和经验复用能力。

核心方法
提出了基于规则的指标体系,通过解框架、执行和反馈控制来评估代理的行为,并通过受控比较评估任务内外的经验复用。
适合谁读
AI研究者、工程师
要解决的问题
现有评估方法忽略了AI代理在长期任务中的行为特征和经验复用能力。
关键实验
对七个前沿模型在36个长期任务上进行了评估。
主要贡献
提供了系统性的评估框架,揭示了现有AI代理更像工程优化器而非完全自主的研究者。
意义与局限
有助于更好地理解AI代理在长期任务中的实际能力,指导未来研究方向;但仅局限于特定任务集上的评估。
领域:cs.AI作者:Yiwei Li、Wanli Yang、Hexiang Tan
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考