MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection· MedFailBench: 医疗AI安全性边界检测的基准
Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels medical AI errors by severity (1--5) and safety gate type (missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, source support gap). The current public release (v0.2.1) contains 44 clinician-reviewed synthetic cases with severity annotations, a live HuggingFace leaderboard preview, a safety gate taxonomy, a clinical severity rubric, and an automated pipeline for archiving model-response screening runs. No patient data, clinical validation claims, or model rankings are included. MedFailBench is r
医疗AI安全边界检测的新基准工具。
- 核心方法
- 构建了一个由临床医生设计的合成基准和错误图谱,标注了医疗AI错误的严重程度和安全门类型。
- 适合谁读
- 研究者、工程师及医疗行业从业人员
- 要解决的问题
- 现有医疗AI基准大多关注模型是否给出正确答案,而忽略了错误的安全边界检测。
- 关键实验
- 当前公开版本(v0.2.1)包含44个临床医生审阅的合成案例,以及HuggingFace领奖台的预览。
- 主要贡献
- 提供了一个开源工具,用于评估和理解医疗AI系统的安全性边界错误。
- 意义与局限
- 有助于提高医疗AI系统的安全性和可靠性,但目前未包含患者数据和临床验证声明。