ai.hackcv
论文精选 65arXiv

Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning· 超越谄媚:大模型道德推理中的结构性抵抗与顺从

Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode. Models must distinguish when to incorporate others' perspectives from when to maintain a well-grounded moral judgment. We study the broader resistance-compliance process governing this distinction. Across three studies, we show that models' judgment revision is structured along three dimensions that parallel classic phenomena in human social psychology: the distance between an incoming view and the model's initial position, the source attribution of that view, and the coalition structure supporting it. Models are generally more receptive to nearby positions, more influenced by views presented as their own

AI 解读论文

揭示大模型在道德推理中如何处理抵抗与顺从的问题

核心方法
通过三个研究,分析模型在判断修订时的结构性抵抗与顺从,重点关注观点距离、来源归属和联盟结构的影响
适合谁读
研究者 / 工程师
要解决的问题
如何使大模型在社会交互中既能学习他人的观点,又能保持自己的道德判断
关键实验
三个研究实验,展示了模型对不同距离、来源和联盟结构的观点的处理方式
主要贡献
提出了大模型在道德推理中的抵抗与顺从的三维结构,并与人类社会心理学现象进行了对比
意义与局限
为构建更符合社会伦理的大模型提供理论基础,但实验设计和数据来源可能有限制
领域:cs.AI作者:Baihui Wang、Bernard Koch
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考