What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models· 合规检测器读取什么?激活探针和防护模型的审计
Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. We show this condition fails across the current class of compliance detectors, a failure we call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe we test, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when that clause is swapped for its permissive counterpart.
审计显示当前合规检测器存在规则盲视问题。
- 核心方法
- 通过对激活探针和防护模型进行实验,研究者测试了删除、排列或替换治理规则对检测准确性的影响。
- 适合谁读
- 研究者 / 工程师 / 法律合规专家
- 要解决的问题
- 当前语言模型的合规检测器在判断时依赖于场景的表面特征而非实际规则,导致规则盲视问题。
- 关键实验
- 实验包括删除、排列或替换治理规则,测试了多种检测器和激活探针,结果显示检测准确性未变。
- 主要贡献
- 揭示了合规检测器的规则盲视问题,并提出需要改进以确保检测器依赖于实际规则。
- 意义与局限
- 该研究指出了当前合规检测器的局限性,对于提高语言模型的法律和审计控制具有重要意义。