ai.hackcv
论文精选 61arXiv

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space· 文本安全之外的物理危险

Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded danger is the same safety problem as ordinary text-level content danger. Through hidden-state direction analysis and random-split null tests, we show that content danger (CD) and physical danger (PD) form separable signals in LLM representations across Qwen2.5-3B/7B/14B/32B, Phi-3.5 and SmolLM2. Building on the CD/PD separability, we propose PRISM, a single-layer L2-regularized logistic probe over full hidden states. PRISM achieves 86.2--87.7\% accuracy on SafeAgentBench with 11.7--13.7\% FPR, while same-scale LLM judges over-block safe tasks at 24.7--39.0\% FPR.

AI 解读论文

研究物理危险与内容危险的区别,提出PRISM探测模型。

核心方法
通过隐藏状态方向分析和随机分割零测试,证明了内容危险和物理危险在LLM表示中形成了可分离的信号,并提出PRISM,一个单层L2正则化逻辑探测模型。
适合谁读
研究者、工程师
要解决的问题
大型语言模型(LLMs)为具身代理提供高级规划时,语言上的无害指令在物理世界中可能变得危险。
关键实验
在Qwen2.5-3B/7B/14B/32B、Phi-3.5和SmolLM2上验证了内容危险和物理危险的可分离性,并测试了PRISM的性能。
主要贡献
PRISM在SafeAgentBench上达到了86.2%-87.7%的准确率,而相同规模的LLM在判定安全任务时的误报率为24.7%-39.0%。
意义与局限
该研究揭示了物理危险与内容危险在大型语言模型中的不同性质,为提高具身代理的安全性提供了新的技术手段。局限性在于PRISM模型的泛化能力和适应性还需进一步验证。
领域:cs.AI作者:Weimeng Wang、Ziqiang Wang、Zihang Zhan
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考