ai.hackcv
论文精选 82arXiv

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents· 工具规格很重要:揭示和缓解 AI 代理的安全风险

AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions. Yet LLMs often become substantially less safe when deployed as agents, and the source of this degradation remains poorly understood. In this paper, we identify schema-formatted tool specifications as a primary source of agent safety degradation and show, through white-box representation analysis, that they weaken the model's internal refusal signals and contribute to unsafe tool execution. Building on this finding, we propose SafeKeep, an inference-time safeguard that decouples safety judgment from tool execution: it assesses requests using flattened textual tool specifications while retaining the original schema-format

领域:cs.AI作者:Minghui Pan、Jiayuxuan Yang、Yuanyuan Yuan
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考