ai.hackcv
论文精选 61arXiv

Evidence-Backed Video Question Answering· 基于证据的视频问答

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations. We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise spatio-temporal evidence: temporal segments and dense, tracked object segmentation masklets. To support this, we introduce ST-Evidence, the first human-verified benchmark for both discriminative and generative pixel-level grounding. Evaluations of state-of-the-art models reveal a critical decoupling between QA accuracy and true visual perception that scaling alone fails to bridge. To address this, we develop scalable, automated generation pipelines to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding. Fine-tuning grounded Video LLMs on this data yields substantial gains over the corresponding size-matched UniPixel baselines (e.g., +27.2 t-mean and +13.8 J&F on a 7B model), establishing a robust baseline for explainable, evidence-backed video understanding. Code and data are available at https://github.com/SalesforceAIResearch/EVQA.

AI 解读论文

提出基于证据的视频问答,增强模型的视觉解释能力。

核心方法
引入E-VQA任务,要求模型同时输出语义答案和精细的空间-时间证据,包括时间片段和密集跟踪的物体分割掩模。通过ST-Evidence数据集支持该任务,该数据集包含人类验证的像素级根据。
适合谁读
研究者 / 工程师
要解决的问题
现有的视频大型语言模型在回答问题时缺乏可验证的视觉依据,且现有解释方法难以捕捉视频中的复杂动态。
关键实验
对现有模型进行了评估,揭示了QA准确性与真实视觉感知之间的关键脱节;并展示了在新数据集上微调后的模型性能显著提升。
主要贡献
创建了ST-Evidence-Instruct数据集,提升了模型在提供精细证据方面的性能,建立了可解释的视频理解基线。
意义与局限
提高了视频问答模型的可解释性和可靠性,对需要精确视觉依据的应用具有重要意义,但生成大量精细证据数据的成本较高。
领域:cs.CV作者:Shijie Wang、Honglu Zhou、Ziyang Wang
相关推荐
论文精选 60已解读

Principia: Relational Physics Tests for Video Models· Principia:视频模型的物理测试

Evaluating physical reasoning in video models is difficult because absolute motion measure…

领域:cs.CV作者:Varun Varma Thozhiyoor、Shivam Tripathi、Venkatesh Babu Radhakrishnan
📎 arXiv🕒 09-04 01:59🔗 arxiv.org

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考