Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?· 桌面Delta基准:计算机使用模型理解桌面GUI转换吗?
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering, and screenshot capture are asynchronous: the next observation may be delayed, occluded, transient, or unrelated, then misread as progress and carried into subsequent planning. We introduce Desktop-Delta Bench (DDB), an offline step-level benchmark with 2,013 human-verified instances from novel, multi-app Linux trajectories across ~15 applications and 50
新基准测试评估计算机使用模型是否能理解桌面GUI转换
- 核心方法
- 创建了包含2013个人工验证实例的离线步级基准DDB,涵盖了多应用程序间的GUI转换
- 适合谁读
- 研究者
- 要解决的问题
- 现有基准测试无法有效评估模型是否能理解由用户操作引起的桌面GUI转换,这对任务完成、错误验证和故障恢复至关重要
- 关键实验
- 未提供
- 主要贡献
- 提供了评估模型对GUI转换理解能力的新标准,促进了此领域的研究进展
- 意义与局限
- 该基准测试有助于识别和改进模型在理解GUI转换时的缺陷,但仅适用于离线评估,受限于特定操作系统的应用范围