ai.hackcv
论文精选 65arXiv

Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs· 临床大模型中证据充足性提示的安全性收益和模型特异性有用性成本

Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency prompt as a test case, we asked whether it reduces unsafe overconfident answers, how far that effect depends on the scoring judge, and what it costs in helpfulness. Methods: In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3) answered a fully paired common panel (1,200 cells) with a standard prompt and the wrapper. The pre-specified endpoint was the paired reduction in unsafe overconfidence scored by the primary judge (GPT-5

AI 解读论文

探讨临床大模型中证据充足性提示的效果及其对不同模型的有用性成本。

核心方法
通过回顾性公共数据基准测试,使用结构化的证据充足性提示,评估四个临床大模型在标准提示和特殊提示下的表现,分析其安全性收益及有用性成本。
适合谁读
研究者、工程师、医疗领域从业人员
要解决的问题
临床大模型在证据不足时容易产生过度自信的回答,影响安全性;研究者尚未明确这种影响是真实的行为改变还是评分者的偏差所致。
关键实验
在Real-POCQi、HealthBench和MedRBench三个数据集上进行了实验,共1,200个样本,评估了四个模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Flash、Grok 4.3)的性能。
主要贡献
揭示了证据充足性提示对不同模型的安全性影响及其特异性有用性成本,有助于理解提示方法对模型行为的影响。
意义与局限
研究结果有助于改进临床大模型的安全性和有用性,但具体影响还需进一步验证;不同模型对提示方法的反应存在差异,这对实际应用提出了挑战。
领域:cs.AI作者:Koyar Afrasyab
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考