ai.hackcv
论文精选 65arXiv

Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis· 多模态大模型的电网诊断任务条件忠实度审计

Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framework in order to conduct task-conditional faithfulness audit. It compares self-reported reliance, intervention-derived behavioral reliance, and preregistered engineering importance. The framework first registers task-specific evidence requirements and compares them with self-reported reliance and behavioral changes under controlled modality ablations. To resolve detected discrepancies, we design an evidence-gated correction and re-audit mechanism that regenerates failed responses under evidence constraints and independently re-ablates them to verify improved groun

AI 解读论文

提出多模态大模型在电网诊断中的任务条件忠实度审计框架。

核心方法
构建一个包含任务特定证据要求注册、自报告依赖性对比、行为依赖性分析及证据约束校正的审计框架。
适合谁读
研究者、工程师
要解决的问题
多模态大语言模型在电网诊断任务中可能未使用适当的证据,导致回答虽准确但缺乏信任度。
关键实验
通过对比自报告依赖性、干预得来的行为依赖性与预注册的工程重要性,验证框架的有效性。
主要贡献
首次提出了一个多模态大模型任务条件下的忠实度审计方法,增强了模型可解释性和可靠性。
意义与局限
提高了多模态大模型在电网诊断任务中的透明度和可靠性,但可能增加模型运行复杂度。
领域:cs.AI作者:Tianqiao Zhao、Meng Yue、Jianhui Wang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考