ai.hackcv
论文精选 65arXiv

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports· 用于肠镜检查的视觉语言基础模型

Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and six multi-centre clinical classification tasks, EndoCLIP outperforms general-purpose and biomedical vision-language encoders in both zero-shot and linear-probe settings. On benign-versus-malignant classification

AI 解读论文

基于大规模肠镜报告开发的视觉语言模型,提高病变检测和报告生成准确性。

核心方法
开发了EndoCLIP模型,使用从280,476份常规肠镜报告中逐步恢复的125,756个病变级别图像-文本对进行训练,以增强图像与文本的对应关系。
适合谁读
医学影像研究人员、AI在医疗领域应用的研发人员、临床医生
要解决的问题
现有的视觉语言模型在肠镜检查中应用较少,主要因为常规报告中的临床发现与其对应图像之间的关联较弱。
关键实验
实验包括病变级别图像-文本检索、结构化报告生成和六个临床分类任务,结果显示EndoCLIP在多个任务中优于现有模型。
主要贡献
EndoCLIP在零样本和线性探针设置下,优于通用和生物医学视觉语言编码器,特别是在病变图像-文本检索、结构化报告生成及六个多中心临床分类任务中表现突出。
意义与局限
该模型的应用有望提高肠镜检查的准确性和效率,但仍需进一步的临床验证。
领域:cs.AI作者:Jia Yu、Yan Zhu、Yili He
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考