ai.hackcv
论文精选 65arXiv

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing· StepGuard: 基于可扩展监督和安全-效用平衡的步级护栏学习

LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed trajectories, leaving pre-execution monitoring of step-level actions underexplored. We propose StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed. To train StepGuard, we introduce StepGen, an automatic data engine that generates safe and unsafe trajectories with the same context but different actions at the risky step. To further reduce over-defense and under-defense, we propose Balance-GRPO, which dynamically balances learning between safe and unsafe actions based o

AI 解读论文

提出一种可扩展监督的步级安全护栏模型,以审计代理轨迹并预检查工具行为。

核心方法
StepGuard通过StepGen自动生成具有相同上下文但不同风险步骤动作的安全和不安全轨迹,结合Balance-GRPO动态平衡学习过程中的安全与不安全行为。
适合谁读
研究者、工程师
要解决的问题
基于LLM的代理使用工具与外部环境互动时存在的安全性风险,如文件修改、信息泄露和未经授权的行为。
关键实验
未提供
主要贡献
提供了一种有效的步级监控机制,能够在行动前预检查工具行为,减少过度防御和防御不足的问题。
意义与局限
该方法可能提高基于LLM代理的安全性,但在实际应用中仍需进一步验证其效果和局限性。
领域:cs.AI作者:Zhijie Zheng、Yu Li、Chen Qian
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考