ai.hackcv
论文精选 61arXiv

Can We Trust Item Response Theory for AI Evaluation?· AI评估中的IRT可靠性

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used

AI 解读论文

评估ILT在AI模型评估中的适用性和可靠性。

核心方法
通过从六个广泛使用的大型语言模型基准中提取项目参数和能力分布,模拟三种常见的IRT模型下的响应矩阵,并比较四种估计工具的性能。
适合谁读
研究者
要解决的问题
论文探讨了标准的IRT估计工具是否适用于AI评估,因为AI基准数据与人类测试数据存在显著差异。
关键实验
使用了六个广泛使用的大型语言模型基准数据进行模拟实验,比较了不同IRT模型和估计工具的性能。
主要贡献
揭示了数据模式不匹配如何影响IRT模型在AI评估中的可靠性,为选择合适的评估工具提供了依据。
意义与局限
为AI评估领域的研究者提供了对IRT模型适用性和局限性的深入理解,有助于改进未来的评估方法。
领域:cs.AI作者:Han Jiang、Sunbeom Kwon、Jinwen Luo
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考