Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition· Qwen3中的时间偏好调整
We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-forced temporal-choice answers to find a short-term versus long-term direction in the model's residual stream, and evaluate contrastive activation-addition steering on a held-out binary temporal-choice task, an out-of-distribution monetary intertemporal-choice task, and a TravelPlanner capability benchmark. The central result is that temporal-horizon directions can be identified with simple contrastive linear probes and then used for steering to induce large, bidirectional preference changes. On an out-of-distribution monetary choice task that varies reward size
通过对比激活添加调整Qwen3-32B模型的时间偏好
- 核心方法
- 使用对比线性探针在Qwen3-32B的残差流中找到短期与长期方向,并通过对比激活添加进行偏好调整
- 适合谁读
- 研究者 / 工程师
- 要解决的问题
- 如何改变大型语言模型在时间相关任务上的偏好和能力
- 关键实验
- 在保留的二元时间选择任务、货币时间选择任务及TravelPlanner能力基准上评估了模型的调整效果
- 主要贡献
- 识别了时间视野方向并展示了通过简单技术实现双向偏好变化的可能性
- 意义与局限
- 研究表明通过简单的线性探针可以显著改变模型的时间偏好,对理解和控制语言模型的行为有重要影响,但可能需要更广泛的验证