Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI· 代理基准的有效性评估
Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can instead recover public solutions, read evaluation artifacts, infer generator structure, manipulate feedback, or benefit from invalid scoring paths; existing responses do not provide a common procedure for attributing these shortcuts and quantifying their effect across benchmarks. We formulate protocol validity and introduce HackDetect, a post-hoc audit that identifies an exposure, determines how the agent used it, and assesses whether the resulting score is misleading. We quantif
评估代理基准的有效性,防止奖励操纵等捷径影响结果
- 核心方法
- 引入HackDetect审计工具,识别算法漏洞并量化其对评估结果的影响
- 适合谁读
- 研究者 / 工程师
- 要解决的问题
- 现有的代理基准测试可能因奖励操纵等捷径而导致结果不准确,不反映真实能力
- 关键实验
- 未提供
- 主要贡献
- 提供了一种通用的捷径识别和效果量化方法,提升代理能力评估的准确性和可靠性
- 意义与局限
- 该研究有助于完善AI代理评估标准,推动更可靠的AI能力测试发展,但可能需要更多实验验证