ai.hackcv
论文精选 65arXiv

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints· 清洁工程,不稳定测量:黑盒LLM观察者可靠性测试失败

Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument'

AI 解读论文

黑盒LLM观察者在共享端点上的可靠性测试失败。

核心方法
研究者进行了两次预先注册的测试活动,固定所有阈值,通过重复请求和次日重放来验证同一请求在不同时间发送给相同模型名称时的一致性。
适合谁读
研究者、工程师
要解决的问题
这篇论文探讨了语言模型观察者在共享端点上进行重复测试时的测量稳定性问题。
关键实验
52,988次审计请求尝试,同一窗口重复排名的相关性为Spearman 0.400,次日重放的相关性为0.78。
主要贡献
揭示了黑盒LLM观察者在共享端点上的测量不稳定性,提出了影响稳定性的三个机制。
意义与局限
该研究强调了在依赖语言模型进行重要决策时需要谨慎处理测量稳定性问题,但未提供具体的解决方案。
领域:cs.AI作者:Haoyaun Zhu、Jie Zhang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考