ai.hackcv
论文精选 60arXiv

SceneActBench: Can Agents Act on the 3D Scenes They See?· SceneActBench:视觉语言模型在3D环境中的行动能力

Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D environment. We evaluate each final output against hidden ground truth with task-specific geometric metrics. SceneActBench comprises five tasks built from 210 source instances, yielding 520 task cases including paired input conditions. Every task runs through one fixed agent loop to keep the comparison

AI 解读论文

探讨视觉语言模型在3D环境中的行动能力

核心方法
提出SceneActBench基准测试,包含五个基于210个源实例的任务,通过统一的代理-环境循环来评估代理基于视觉条件的行动
适合谁读
研究者、工程师
要解决的问题
现有3D基准测试未能充分评估视觉语言模型在处理完整多物体3D场景时的行动能力
关键实验
包含了520个任务案例,使用任务特定的几何度量标准对最终输出进行评估
主要贡献
提供了一个全面评估视觉语言代理在3D场景中执行任务的基准测试框架
意义与局限
意义在于推动3D环境中的视觉语言模型研究,影响可能包括提高机器人在真实世界的行动能力,局限在于目前的任务设置和评估标准可能不涵盖所有现实情况
领域:cs.AI作者:Yifei Zhao、Xiangxin Zhou、Wenhao Yang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考