ai.hackcv
论文精选 65arXiv

Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement· 基于超图的多模态检索与生成框架

Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of traditional simple graphs, which fails to capture the intricate, high-order correlations among heterogeneous entities, such as the N-ary relationships between a visual chart, its scattered textual descriptions, and underlying numerical data. Furthermore, existing refinement strategies often rely on exhaustive, full-page reconstruction to align cross-modal information, leading to prohibitive computational redundancy and the introduction of contextual noise in long-form document processing. In this paper, we propose Hyper-M2RAG, a novel framework that redefines multimodal document retrieval through High-order Hypergraph Representation Learning. We first formalize

AI 解读论文

提出基于超图的多模态检索与生成框架,解决传统图模型无法捕捉复杂高阶关联的问题。

核心方法
通过高阶超图表示学习重新定义多模态文档检索框架,利用超图结构捕捉图表、文本描述和数值数据间的N元关系;采用增量优化策略更新文档,降低计算冗余。
适合谁读
研究者、工程师
要解决的问题
传统简单图模型在多模态文档检索中只能表示二元连接,无法捕捉不同实体间的复杂高阶关联;现有优化策略在处理长文档时计算冗余且引入上下文噪声。
关键实验
在多个多模态检索与生成数据集上进行实验,验证了Hyper-M2RAG在跨模态信息检索、长文档生成等任务上的优越性能,具体指标相比于现有方法有明显提升。
主要贡献
引入超图结构用于多模态信息的高阶关联建模;提出增量优化方法提升长文档处理效率;在多项任务上实现性能显著提升。
意义与局限
为多模态信息处理提供了新的视角和工具,有助于提高跨模态数据的精准匹配和高效生成;在处理复杂高阶关联方面具有潜在优势,但超图的构建与优化可能带来新的技术挑战。
领域:cs.AI作者:Shenao Chen、Yidan Xu、Xiangmin Han
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考