ai.hackcv
论文精选 60arXiv

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework for evaluating the reliability of concept-based explanations. Rather than replacing SMILE, ConceptSMILE extends its perturbation-based logic from feature- or region-level attribution to the auditing of human-understandable concept explanations. The framework perturbs input regions, measures concept-response shifts, applies locality weighting, and fits an XGBoost surrogate to approximate local concept behaviour. Reliability is assessed through attribution accuracy, surrogate fidelity, faithfulness, stability, and consistency. We evaluate ConceptSMILE on retinal fundus images by comparing MedSAM-derived visual concepts with VLM-based semantic concepts. Results show that reliability varies across concepts and pathways: MedSAM achieves stronger spatial attribution and the highest surrogate fidelity ($R^2 = 0.8503$, $R_w^2 = 0.8465$), while the VLM pathway shows stronger vessel faithfulness and stronger stability under selected artefact conditions. ConceptSMILE provides an independent audit layer for evaluating the trustworthiness of concept-based XAI.

AI 解读论文

一种评估概念解释可靠性的框架,适用于医学图像领域。

核心方法
提出ConceptSMILE,基于扰动的模型无关审计框架,通过扰动输入区域、测量概念响应变化、应用局部权重和拟合XGBoost代理模型来评估概念解释的可靠性。
适合谁读
研究者、工程师、医疗AI应用开发者
要解决的问题
概念解释在解释AI模型时虽易懂,但其可靠性难以保证。
关键实验
使用视网膜眼底图像,比较了MedSAM派生的视觉概念与基于VLM的语义概念的可靠性。
主要贡献
为概念解释提供了一个独立的审计层,能够多维度评估其可靠性,如归因准确性、代理保真度、可信度、稳定性和一致性。
意义与局限
有助于提高概念解释的透明度和可信度,促进医疗领域AI系统的安全应用;局限性在于目前仅在特定数据集上进行了评估。
领域:cs.AI作者:Mohadeseh Mollapour、Koorosh Aslansefat、Zeinab Dehghani
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考