Motion-Conditioned Multi-View Fusion for Myocardial Infarction Localization from Echocardiography· 心肌梗塞多视角定位的运动条件融合方法
Myocardial infarction (MI) remains a leading cause of mortality worldwide. Echocardiography (Echo) is a widely available modality for MI assessment, where regional wall motion abnormality is a key indicator. Prior learning based methods for myocardial motion analysis often use handcrafted descriptors or densely supervised estimation, but the need for extensive annotation limits applicability. Foundation models have recently improved vision-based Echo analysis; however, most methods operate on single views and segment-level localization remains unreliable under view-dependent ambiguity, especially in apical views. To address this, we propose MCF-Net, a novel motion-guided multi-view fusion framework that fuses myocardial motion cues with foundation model representations to localize infarction. Visual features are extracted using EchoPrime, a pretrained Echo foundation model shared across dual views. Cardiac motion is modeled with extremely sparse supervision: a single annotated template frame is transferred across videos to initialize point tracking, avoiding dense labels. Motion-derived segment-aware soft masks provide coarse spatial priors that selectively enhance features for challenging myocardial segments. A motion-conditioned fusion mechanism then integrates motion and vision across views, refining predictions without overriding strong appearance cues. On segment-level MI localization, MCF-Net achieves 72.4\% F1 and 84.9\% accuracy, outperforming state-of-the-art motion-only, vision-only, and fusion baselines.
心肌梗塞多视角定位的新方法
- 核心方法
- 提出MCF-Net,利用预训练Echo模型提取多视角视觉特征,并结合稀疏监督的运动模型进行融合定位
- 适合谁读
- 医学影像研究者、心脏病学专家、AI算法工程师
- 要解决的问题
- 心肌梗塞区域定位不准确,尤其在心尖视角下
- 关键实验
- 在段级心肌梗塞定位任务中,F1得分72.4%,准确率84.9%,优于现有方法
- 主要贡献
- 实现了更高的段级心肌梗塞定位精度,减少依赖密集注释
- 意义与局限
- 提高心肌梗塞诊断的准确性和可靠性,为临床应用提供新工具;但可能受限于不同心脏类型的数据多样性