ai.hackcv
论文精选 65arXiv

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context· 稳健性的错觉:总体准确率掩盖了任务无关上下文下的预测变化

As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy. This aggregate stability, however, masks significant per-example instability. Even semantically meaningless pseudo-words, formed by randomly combining characters, can markedly shift model predictions on a small fraction of examples, degrading performance on some while improving it on others. This two-sided effect holds consistently across a wide range of models and datasets, yet the affected examples are largely model-specific. We further show that this instability is modulated by context type, context length, test-time compute, and model development stage. Together, our findings reveal context-induced tail risks concealed by aggregate accuracy, motivating per-example reliability evaluation of language models.

AI 解读论文

大语言模型总体准确率掩盖了个别预测的不稳定性。

核心方法
通过在基准测试问题前附加与任务无关的上下文(包括无意义的伪词),研究者系统地分析了不同模型和数据集上的预测变化,考察了上下文类型、长度、测试时的计算资源以及模型开发阶段对这种不稳定性的影响。
适合谁读
研究者、工程师、AI 系统评估专家
要解决的问题
这篇论文探讨了大语言模型在任务无关上下文中的表现是否真的稳健,即其在包含部分无关信息的输入下是否能保持预测的一致性。
关键实验
在多个模型和数据集上进行了实验,添加了不同类型的上下文(包括无意义伪词),并考察了不同参数对预测变化的影响。
主要贡献
揭示了大语言模型在任务无关上下文下存在显著的个体样本不稳定性,即使总体准确率较高;强调了在评估语言模型时考虑每样本可靠性的必要性。
意义与局限
该研究指出了仅依赖总体准确率评估模型稳健性的局限性,强调了对个别样本进行细致分析的重要性,有助于提高未来模型的可靠性和安全性。
领域:cs.CL作者:Yanzhe Zhang、Sanmi Koyejo、Diyi Yang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考