OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models· 跨平台计算机使用奖励模型标准化评估
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, then rig
介绍OSReward,用于评估视觉-语言模型在计算机使用代理轨迹上的可靠性。
- 核心方法
- 构建OSReward基准测试,该测试由多样化的代理执行经人类验证的指令产生的轨迹组成,用于系统性评估视觉-语言模型的判断能力。
- 适合谁读
- 研究者 / 工程师
- 要解决的问题
- 现有视觉-语言模型在评估计算机使用代理执行任务轨迹方面缺乏可靠性验证。
- 关键实验
- 展示了OSReward在多种现有视觉-语言模型上的评估结果,并与人类判断进行了对比。
- 主要贡献
- 提供了首个评估视觉-语言模型作为计算机使用代理轨迹评判者的可靠性的基准测试OSReward。
- 意义与局限
- 该研究为视觉-语言模型在代理评估领域的应用提供了重要的参考依据,但仍然存在模型偏差和人类判断复杂性的问题。