ai.hackcv
论文精选 65arXiv

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?· 桌面Delta基准:计算机使用模型理解桌面GUI转换吗?

Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering, and screenshot capture are asynchronous: the next observation may be delayed, occluded, transient, or unrelated, then misread as progress and carried into subsequent planning. We introduce Desktop-Delta Bench (DDB), an offline step-level benchmark with 2,013 human-verified instances from novel, multi-app Linux trajectories across ~15 applications and 50

AI 解读论文

新基准测试评估计算机使用模型是否能理解桌面GUI转换

核心方法
创建了包含2013个人工验证实例的离线步级基准DDB,涵盖了多应用程序间的GUI转换
适合谁读
研究者
要解决的问题
现有基准测试无法有效评估模型是否能理解由用户操作引起的桌面GUI转换,这对任务完成、错误验证和故障恢复至关重要
关键实验
未提供
主要贡献
提供了评估模型对GUI转换理解能力的新标准,促进了此领域的研究进展
意义与局限
该基准测试有助于识别和改进模型在理解GUI转换时的缺陷,但仅适用于离线评估,受限于特定操作系统的应用范围
领域:cs.AI作者:Abhishek Pillai、Samir Kumar Nayak、Yuan Chen
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考