ai.hackcv
论文精选 61arXiv

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy· 多模态大模型科学可视化理解能力评估

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whe

AI 解读论文

评估多模态大模型在科学可视化理解上的表现

核心方法
使用包含49个项目、18个科学可视化和11种任务类型的标准化评估测试,对六个多模态大模型进行评估
适合谁读
研究者、工程师
要解决的问题
现有评估方法主要集中在图表上,缺乏对科学可视化理解的全面评估
关键实验
评估了三个闭源和三个开源模型,使用485名人类参与者的数据进行对比
主要贡献
揭示了当前多模态大模型在科学可视化理解上的能力和局限
意义与局限
为多模态大模型在科学可视化领域的应用提供了基准,指出了改进方向
领域:cs.AI作者:Patrick Phuoc Do、Chau M. Ta、Chaoli Wang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考