ai.hackcv
论文精选 65arXiv

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation· Messier:跨基准代理评估的高分辨率语料库

Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of the empirical record incomparable. We introduce Messier, a unified corpus of 957,253 records that span 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. Messier consolidates public benchmark scores and supplements them with five-agent runs across six underrepresented professional and scientific domains, including a recent legal benchmark. Each record is standardized by model, scaffold, environment, task, verifier, and aggregation rule, with SOC/NAICS classifications for occupational and industry analysis. Using this corpus, we show frontier progres

AI 解读论文

跨基准评估 AI 代理的高分辨率语料库 Messier 介绍。

核心方法
构建了一个统一的高分辨率语料库 Messier,包含 30 个基准、714 个代理、11,891 个任务和 74,205 个验证器的 957,253 条记录,通过 SOC/NAICS 分类系统进行标准化。
适合谁读
研究者
要解决的问题
现有 AI 代理评估方法受制于任务、支架、验证器和评分规则的碎片化,导致评估结果难以比较。
关键实验
未提供
主要贡献
提供了跨越多个专业和科学领域的标准化评估记录,增强了 AI 代理评估的可比性和全面性。
意义与局限
该语料库为 AI 代理的跨基准评估提供了重要资源,但尚未包含实验验证其有效性和应用效果。
领域:cs.AI作者:Stefan Krsteski、Charlotte Meyer、Guillaume Allegre
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考