SciMIF: Understanding Multimodal Instruction Following in Scientific Domains· SciMIF: 理解科学领域多模态指令遵循
Understanding instruction-following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development of scientific fields. In this work, we introduce SciMIF, a novel benchmark designed to evaluate the capability of MLLMs in following complex scientific instructions. Specifically, based on an extensive analysis of 22 distinct tasks across 5 representative scientific disciplines, we propose a comprehensive taxonomy comprising 10 constraint groups that captures both general functional requirements and discipline-specific characteristics. Guided by this taxonomy, we develop a high-fidelity instruction injection pipeline to systematically augment existing scientific datasets. We conduct comprehensive experiments on mu
提出科学领域多模态指令遵循的评估基准SciMIF,促进科学发展的多模态大模型应用。
- 核心方法
- 基于5个代表性科学领域的22项不同任务,提出包含10个约束组的综合分类法,并开发高保真指令注入管道以增强现有科学数据集。
- 适合谁读
- 研究者、工程师
- 要解决的问题
- 评估多模态大语言模型在科学领域遵循复杂指令的能力。
- 关键实验
- 进行了全面的实验,但具体实验结果和数据未完全提供。
- 主要贡献
- 建立了第一个针对科学领域多模态指令遵循能力的评估基准SciMIF,推动了多模态大模型在科学领域的应用研究。
- 意义与局限
- 为科学领域中的多模态大语言模型提供了标准化评估,有助于模型的进一步优化和实际应用,但也需注意基准覆盖的领域和任务可能有限。