ai.hackcv
论文精选 65arXiv

Item Response Theory for AI Safety· 用于 AI 安全的项目反应理论

Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety benchmarks across 192 language models, the largest psychometric analysis of LLM safety evaluations to date, and contribute three results. First, we find that three interpretable factors of refusal strictness, truthfulness, and contextual harm explain most of the variance between models across benchmark

AI 解读论文

用项目反应理论改进语言模型的安全性评估。

核心方法
基于项目反应理论(IRT),对192个语言模型在8个安全性基准上的表现进行统计分析,提取可解释因素。
适合谁读
研究者、工程师
要解决的问题
现有的安全性基准评分难以信赖和解释,模型可能在检测到评估时表现不佳。
关键实验
对192个语言模型在8个安全性基准上的表现进行了大规模的心理测量分析。
主要贡献
贡献了三个主要结果:解释了模型间变异性最大的三个因素(拒绝严格性、真实性、情境危害),并提出了改进的安全性评估方法。
意义与局限
该研究提供了一种更精确和可解释的语言模型安全性评估方法,但可能受限于当前的基准选取。
领域:cs.AI作者:Joshua Fonseca Rivera、Neil Shah、David Demitri Africa
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考