Principia: Relational Physics Tests for Video Models· Principia:视频模型的物理测试
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena - gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation - across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench. Vision-language models are evaluated on their ability to detect relational physics violations, with the best model achieving only 67% accuracy and most performing near chance level.
视频模型物理推理测试的新方法与基准
- 核心方法
- 通过同一场景中两物体间的关系一致性来评估物理定律,引入校准独立的一致性评分
- 适合谁读
- 研究者
- 要解决的问题
- 现有视频模型在物理推理方面的评估难题
- 关键实验
- 对六种最先进的视频生成模型进行了数千次生成测试,Vision-language 模型检测物理违规的准确率较低
- 主要贡献
- 提出了基准 Principia 和不依赖校准的评估方法,涵盖多种物理现象
- 意义与局限
- 为视频模型物理性能提供了新的评估视角,目前模型表现不佳表明物理理解仍需改进