SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?· SAEScientist-Bench:AI代理能否自主进行SAE可解释性研究
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Bench to evaluate whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery. Given a target concept, an agent designs contrastive probes and navigates a Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT to discover the optimal feature, evaluated against c
AI代理能否自主进行SAE模型的可解释性研究
- 核心方法
- 通过引入SAEScientist-Bench评估AI代理使用SAE工具自主进行机制发现的能力,以目标概念设计对比探针并导航特征字典
- 适合谁读
- 研究者、工程师
- 要解决的问题
- 当前的递归自我改进研究主要集中在自动化模型训练,缺乏对手模型学习内容的后处理监测和审计
- 关键实验
- 实验设计包括代理在给定目标概念下通过对比探针发现最优特征的性能评估
- 主要贡献
- 提出了SAEScientist-Bench,一个评估AI代理在SAE可解释性研究中表现的新基准
- 意义与局限
- 为促进可靠自主发展的AI系统提供了重要工具,但当前实现可能受限于特定领域和代理的能力