Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model· 材料科学机制在开放权重语言模型中的表征
Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both
探讨材料科学机制在开放权重语言模型中的表征形式及其验证方法。
- 核心方法
- 通过结合匹配的直接和雅可比词汇读出、无选项状态几何、60定律反事实基准和因果干预,分离并验证材料科学机制在模型中的三种表现形式:概念、构成方向和内部表征。
- 适合谁读
- 材料科学和AI研究者
- 要解决的问题
- 现有大型语言模型在回答科学问题时,无法明确其是否真正理解或使用了基础物理知识。
- 关键实验
- 在50个保留的材料描述中,独立拟合的三个雅可比透镜成功再现了概念排名,且无需特定目标词汇集。
- 主要贡献
- 首次系统性地证明了材料科学机制在开放权重语言模型中的存在形式,并提供了验证方法。
- 意义与局限
- 该研究为理解语言模型如何处理专业科学知识提供了新视角,但局限于特定模型和方法,需进一步扩展。