Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark, and have not focused on understanding the model's error types. Here, we introduce the Complex Social Behavior (CSB) dataset, containing 100 images depicting complex social interactions/behaviors. We analyze the progression of scene descriptions over a decade (2017-2025) of VLMs (four pre-Multimodal Large Language Models, MLLMs, and five MLLMs). We evaluate the accuracy of the models and 20 human descriptions relative to a gold standard on the CSB dataset and on a sample from MS-COCO. We analyzed five visual-cognitive error types: object detection, recognition, hallucination, scene understanding, and spatial dependence. The CSB dataset showed a more pronounced improvement than MS-COCO in scene description accuracy, with pre-MLLMs achieving much lower accuracy than the bottom-ranked human descriptions and MLLMs attaining accuracies similar to the top-ranked human descriptions. We show that MLLMs have eliminated the gap in scene description accuracy between simpler MS-COCO scenes and scenes depicting complex behaviors (CSB). MLLMs have almost eliminated all error types in our tested datasets, except for occasionally relying on different image regions for scene descriptions than humans do (spatial dependence error). We also show that detection, recognition, and hallucination errors have the highest impact on scene description accuracy. Together, our findings provide a more thorough evaluation of how visual language models have advanced over the last decade.
过去十年视觉语言模型在复杂社交行为图像描述上的准确性和视觉认知错误分析
- 核心方法
- 引入 Complex Social Behavior (CSB) 数据集,评估 2017-2025 年间 VLMs 和 MLLMs 的场景描述准确率,并分析了五种视觉认知错误类型
- 适合谁读
- 研究者 / 工程师 / 产品经理
- 要解决的问题
- 论文解决了视觉语言模型在复杂社交行为场景描述上的准确性和错误类型分析问题
- 关键实验
- 在 CSB 数据集和 MS-COCO 样本上评估了九个模型和 20 个来自人类的描述,分析了多种视觉认知错误类型
- 主要贡献
- 展示了 MLLMs 在复杂社交行为场景描述上的显著进步,几乎消除了与人类描述的差距
- 意义与局限
- 研究提供了对视觉语言模型十年进展的深入理解,揭示了 MLLMs 在处理复杂社交行为场景时的能力和局限性