Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication· 超越记录:检测隐秘多代理通信
Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, creating opportunities for covert harmful coordination. We introduce Verifiable Latent Alignments (VLA), an activation-aware framework for monitoring and steering these private communication channels. For every monitored decision, VLA links the private latent-state record and channel status to the resulting public action using a shared event identifier, enabling matched causal analysis. Our first contribution is a neutral-only three-layer monitor combining representation anomaly detection, counterfactual action-distribution influence, and sparse-autoencoder interpretation support. Our second contribution is a steerability framework spanning black-box behavioral instructions and
提出框架检测和控制隐秘多代理通信
- 核心方法
- 引入可验证隐态对齐(VLA)框架,结合表征异常检测、反事实行为影响分析和稀疏自动编码器解释支持
- 适合谁读
- 研究者、工程师
- 要解决的问题
- 解决语言模型代理通过隐藏状态进行隐秘和有害协调的问题
- 关键实验
- 未提供
- 主要贡献
- 提出中立的三层监测模型和黑盒行为指令的可控制框架
- 意义与局限
- 为监测和控制隐秘代理通信提供了新工具,但需要进一步实验验证其效果和局限性