ai.hackcv
论文精选 65arXiv

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models· 跨平台计算机使用奖励模型标准化评估

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, then rig

AI 解读论文

介绍OSReward,用于评估视觉-语言模型在计算机使用代理轨迹上的可靠性。

核心方法
构建OSReward基准测试,该测试由多样化的代理执行经人类验证的指令产生的轨迹组成,用于系统性评估视觉-语言模型的判断能力。
适合谁读
研究者 / 工程师
要解决的问题
现有视觉-语言模型在评估计算机使用代理执行任务轨迹方面缺乏可靠性验证。
关键实验
展示了OSReward在多种现有视觉-语言模型上的评估结果,并与人类判断进行了对比。
主要贡献
提供了首个评估视觉-语言模型作为计算机使用代理轨迹评判者的可靠性的基准测试OSReward。
意义与局限
该研究为视觉-语言模型在代理评估领域的应用提供了重要的参考依据,但仍然存在模型偏差和人类判断复杂性的问题。
领域:cs.AI作者:Qiushi Sun、Kanzhi Cheng、Yian Wang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考