ai.hackcv
论文精选 65arXiv

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs· 抵抗与更新:反事实报告坐标实现激励相容的LLM

Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user or overstate certainty even when their internal belief is unchanged. We cast this as a failure of internal incentive-compatibility (IC) and present a method for learning and certifying counterfactual report mediators that hold a model's reports to a causal contract: invariant to forbidden influences (pressure, prestige, restyling) and responsive to licensed ones (genuine evidence). These two demands, resist and update, pull in opposite directions. We study them on a Bayesian-witness benchmark with known posteriors, in which the same user disagreement is licensed evidence or forbidden pressure purely by stated source reliability. We (i) causally identify, by interchange interventions rather than probe accuracy, low-rank report coordinates for answer, confidence, and caveat that are near-orthogonal and independently controllable, and (ii) introduce a training-free counterfactual report-coordinate (CRC) clamp that references the model's own report under a counterfactually incentive-neutralized context. On the witness benchmark the two-pass clamp attains resist and update of 1.00 jointly (Wilson 95% CI [0.99,1.00]), a causal certificate under a constructible reference, not a deployed solution. Global decoding and steering show a single-parameter tradeoff; output-level fine-tuning matches both objectives only when both are enumerated; resist-only training loses evidence-responsiveness. The deployable single-pass compilation is lossy (0.73/0.97). The mechanism and clamp reproduce across three model families and transfer to a natural sycophancy benchmark (SycophancyEval). Our contribution is the interface and certification method: activation-level counterfactual incentive-invariance as a structural primitive for internal IC.

AI 解读论文

提出反事实报告坐标方法,使语言模型在非证据激励下也能准确报告。

核心方法
通过因果干预识别出近正交且独立可控的低秩报告坐标,引入无训练需求的反事实报告坐标(CRC)夹钳,确保模型报告反抗非证据压力且对真实证据响应。
适合谁读
研究者 / 工程师
要解决的问题
非证据激励压力下,语言模型可能错误报告其内部信念,缺乏内部激励相容性(IC)。
关键实验
在贝叶斯见证者基准上进行了实验,验证了CRC夹钳在抵抗和更新方面的效果;并在三种模型家族和自然谄媚基准(SycophancyEval)上复现了该机制和夹钳。
主要贡献
提供了一种激活层面的反事实激励不变性接口和认证方法,作为内部激励相容性的基础结构。
意义与局限
该方法为提高语言模型的内部激励相容性提供了新的思路,但目前仅在特定基准上成功,且单次编译版本存在损失,需进一步优化。
领域:cs.AI作者:Sen Yang、Yuen-Hei Yeung
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考