ai.hackcv
论文精选 65arXiv

Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification· 更智能验证,更进一步进化

Agent harnesses shape how language-model agents use instructions, tools, and runtime components, but adapting these harnesses requires costly verification. Existing propose-and-verify methods typically score every candidate on a fixed task set, wasting rollouts on unrelated behaviors and allowing aggregate scores to obscure specific regressions. We introduce HarnessLens, a budget-aware framework for automated harness evolution. HarnessLens jointly explores the task space and user-configurable components, derives candidate modifications from execution trajectories, and selectively verifies each candidate on behavior-relevant tasks using an attributable-evidence gate. Across three agent harnesses and four benchmarks, HarnessLens improves average held-out performance by 7.6-13.6% while consum

AI 解读论文

通过行为感知验证优化代理框架,提高性能并减少资源浪费。

核心方法
提出HarnessLens框架,结合任务空间探索和用户可配置组件,从执行轨迹中推导候选修改,并通过可归因证据门选择性地验证行为相关的任务。
适合谁读
研究者、工程师
要解决的问题
现有的代理框架进化方法在验证阶段存在资源浪费和无法精准识别性能退化的问题。
关键实验
在三个代理框架和四个基准上进行了实验,验证了HarnessLens的有效性。
主要贡献
改善了三个代理框架在四个基准上的平均性能,同时减少了验证过程中的资源消耗。
意义与局限
该方法提高了代理框架的进化效率,对资源受限的环境特别有意义,但可能需要更多的计算资源来支持初期的任务空间探索。
领域:cs.AI作者:Jinghan Xu、Yikai Zhang、Aili Chen
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考