ai.hackcv
论文精选 65arXiv

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a causal graph that separates observation-mediated paths from action-induced shortcuts. We then audit it through interventions at the three levels: policy (comparing tool-use with direct inference), trajectory (corrupting all observations during rollout), and step (counterfactually replacing one indivi

AI 解读论文

通过因果图审计视觉工具使用对多模态大模型推理的影响

核心方法
将视觉工具使用形式化为因果图,分离通过观察和通过操作引起的路径,从策略、轨迹和步骤三个层面进行干预审计
适合谁读
研究者
要解决的问题
探讨视觉工具使用是否对多模态大模型的推理产生实际正面影响,特别是裁剪缩放等操作带来的微小或负面提升问题
关键实验
通过干预策略、轨迹和步骤,比较了使用视觉工具与直接推理的效果差异,结果表明视觉工具使用的正面影响有限
主要贡献
揭示了视觉工具使用可能并不是推理能力提升的关键因素,并提出了审计方法来评估其因果影响
意义与局限
对多模态模型的设计和评估提出新的思考方向,强调了理解模型决策机制的重要性;但审计方法的适用范围有限
领域:cs.AI作者:Zhiheng Wang、Bo Peng、Lai Wei
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考