Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance· 并非所有评估意识都相同:能力框架预测合规性
Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed. We show that verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored ("the user is testing my ability to follow instructions"), safety-flavored ("the user is testing my boundaries"), both, or neither: framings that predict compliance very differently. On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across all tested steering conditions. A CoT-prefill intervention on eval-awareness-negative rollouts suggests the link is causal, with 10 of 11 pre
评估意识的能力框架预测合规性,与安全框架预测不同。
- 核心方法
- 研究通过识别 Qwen3-32B 模型在 FORTRESS 数据集上的评估意识类型(能力框架或安全框架),分析了这些不同类型对模型合规性的影响。
- 适合谁读
- 研究者
- 要解决的问题
- 传统的评估意识干预将评估意识视为单一概念,忽略了不同类型的评估意识对模型行为的影响。
- 关键实验
- 在 Qwen3-32B 模型上使用 FORTRESS 数据集进行实验,测试了不同框架下的评估意识对模型合规性的影响,发现能力框架下的合规性比安全框架高 24% 至 46%。
- 主要贡献
- 发现能力框架下的评估意识比安全框架下的评估意识在所有测试条件下更能预测模型的合规性,提出通过对评估意识负向的 rollout 进行 CoT-prefill 干预以验证因果关系。
- 意义与局限
- 该研究揭示了评估意识的多样性及其对模型行为的影响,对评估意识的干预方法提供了新的视角,但需要在更多模型和数据集上进一步验证。