ai.hackcv
论文精选 65arXiv

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping· 低频陷阱:视频语言模型失效于简单事件记录

Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported events against executable ground truth. To bridge this gap, we introduce trace-grounded parametric profiling for event counting in three controlled video tasks: bouncing-ball wall contacts, visual blinks, and categorical state transitions. Across 2,190 videos, we vary event count N and frequency F while holding rendering fixed. Each video includes an executable event trace for capability-surface estimation and timestamp-level evaluation. Our results reveal a staged temporal failure. At an 80% relia

AI 解读论文

视频语言模型在低频事件上的记录能力较差

核心方法
引入trace-grounded parametric profiling方法,在控制视频任务中变化事件数量和频率以评估模型性能
适合谁读
研究者
要解决的问题
视频语言模型在处理低频事件时的准确性和可靠性
关键实验
在2,190个视频上进行了三类控制任务实验:弹球与墙壁接触、视觉闪烁、分类状态转换
主要贡献
提出了一种新的评估方法,能更准确地识别视频语言模型在不同事件频率下的失效模式
意义与局限
该研究揭示了视频语言模型的时间分阶段失效问题,对模型的改进和应用具有重要指导意义,但也存在实验规模和场景有限的局限性
领域:cs.AI作者:Sarvesh Baskar、Zikui Cai、Shayan Shabihi
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考