Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought
Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the visible chain plays that role is rarely tested. General-domain CoT-faithfulness probes ignore clinical cost, and medical LLM evaluations treat the chain as a black box. We close this gap with a medical perturbation audit: a 30-operator battery edits both the chain and the question with clinically motivated operators (severity reversal, negation flip, demographic swap, evidence ablation), paired with a chain-update times answer-flip joint analysis that classifies each model by its failure mode. Applied to 14 LLMs on four medical QA benchmarks, three independent tests converge: the Chain-Decoupling Rate (CDR; chain does not register the edit and the answer does not flip) is 72.9% panel-wide on
研究通过扰动审计评估了14种LLM在医学问答中的推理链可靠性。
- 核心方法
- 设计了包含30种操作符的医学扰动审计方法,编辑问题和推理链,通过链更新与答案改变联合分析来分类模型的失败模式。
- 适合谁读
- 研究者、工程师
- 要解决的问题
- 现有的医学LLM评估方法没有充分测试推理链的可见性是否真正起到了医疗推理的作用。
- 关键实验
- 对14种LLM进行测试,使用四个医学问答基准,结果显示整体CDR为72.9%。
- 主要贡献
- 介绍了医学扰动审计框架,量化了医学LLM的推理链解耦率(CDR),揭示了现有LLM在医学领域的推理能力缺陷。
- 意义与局限
- 该研究强调了医学领域中LLM推理链透明性和可靠性的必要性,为改进医学AI模型提供了方向。但仅测试了现有的LLM,未涉及模型改进的具体方法。