From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of the objective or strategy producing it. We test these distinctions in two open-weight model families. Across controlled guessing-game and stock-trading experiments, we find that deceptive-looking behavior can arise without the corresponding proposed mechanism, while other interventions provide direct
通过因果框架区分语言模型欺骗行为和机制
- 核心方法
- 引入因果分类法,区分先验承诺与回顾报告、模型偏好与实际输出、虚假偏好与误导收益敏感性、欺骗行为与产生该行为的策略
- 适合谁读
- 研究者
- 要解决的问题
- 解决语言模型欺骗研究中过度拟人化的问题,澄清行为与机制的区别
- 关键实验
- 在两个开放权重模型家族中进行了受控猜测游戏和股票交易实验
- 主要贡献
- 提供了一种新的分类框架,帮助更准确地理解语言模型中的欺骗现象
- 意义与局限
- 有助于避免误解语言模型行为,促进更严谨的欺骗研究;但实验范围有限,可能不完全代表所有模型