ai.hackcv
论文精选 65arXiv

Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication· 超越记录:检测隐秘多代理通信

Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, creating opportunities for covert harmful coordination. We introduce Verifiable Latent Alignments (VLA), an activation-aware framework for monitoring and steering these private communication channels. For every monitored decision, VLA links the private latent-state record and channel status to the resulting public action using a shared event identifier, enabling matched causal analysis. Our first contribution is a neutral-only three-layer monitor combining representation anomaly detection, counterfactual action-distribution influence, and sparse-autoencoder interpretation support. Our second contribution is a steerability framework spanning black-box behavioral instructions and

AI 解读论文

提出框架检测和控制隐秘多代理通信

核心方法
引入可验证隐态对齐(VLA)框架,结合表征异常检测、反事实行为影响分析和稀疏自动编码器解释支持
适合谁读
研究者、工程师
要解决的问题
解决语言模型代理通过隐藏状态进行隐秘和有害协调的问题
关键实验
未提供
主要贡献
提出中立的三层监测模型和黑盒行为指令的可控制框架
意义与局限
为监测和控制隐秘代理通信提供了新工具,但需要进一步实验验证其效果和局限性
领域:cs.AI作者:Ramneet Kaur、Pradyumna Chari、Ramesh Raskar
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考