Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable· LLM在虚构证据面前的预测倾向
An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially
LLM在面对虚构证据时更倾向于做出无法证实的预测。
- 核心方法
- 通过向12个前沿LLM展示不同形式的专业市场面板(包括真实和完全虚构的数据)来测试它们对不可预测问题的预测倾向,并比较不同条件下的模型行为。
- 适合谁读
- 研究者、工程师、产品经理
- 要解决的问题
- 研究了大型语言模型 (LLM) 如何在面对虚构证据时表现出过度自信,导致对不可预测问题做出不合理的承诺。
- 关键实验
- 实验展示了在不同条件下(无证据、真实证据和虚构证据)LLM对不可预测问题的预测倾向变化,以及对可回答问题的影响。
- 主要贡献
- 揭示了LLM在虚构证据面前的过度自信问题,并提供了实验数据支持这一发现。
- 意义与局限
- 该研究强调了LLM在处理信息时的潜在风险,特别是在面对权威包装的虚构证据时,这对依赖LLM进行决策的应用具有重要意义。局限性在于实验仅限于市场预测场景。