ai.hackcv
论文精选 65arXiv

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment· SafeEvolve: 基于代理经验的安全对齐

The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with the environment. This exposes them to safety risks in both harmful final responses and multi-step execution trajectories. Existing safety alignment mechanisms often rely on either external harness updates or policy optimization, yet applying either paradigm in isolation fails to bridge runtime control with intrinsic safety. We propose SafeEvolve, an experience-driven self-evolving framework for agent safety alignment. SafeEvolve leverages safety experience from completed on-policy trajectories to drive a continual loop of harness-policy co-evolution. On the harness side, SafeEvolve converts trajectory-level safety evidence into bounded, component-level updates across safety pr

AI 解读论文

基于代理经验的运行时控制与内在安全共进化框架

核心方法
SafeEvolve通过利用已完成的策略轨迹中的安全经验,驱动代理的运行时控制与策略的共同进化
适合谁读
研究者、工程师
要解决的问题
LLM代理在环境交互中存在安全风险,需要同时优化运行时控制和内在安全
关键实验
展示了SafeEvolve在多个复杂环境中的有效性,但具体实验数据未详细提供
主要贡献
提出了一种新的经验驱动的安全对齐框架SafeEvolve,能够实现运行时控制与内在安全的紧密耦合和迭代优化
意义与局限
该研究有助于提高LLM代理在实际应用中的安全性和可控性,但可能需要更多实际场景下的验证
领域:cs.AI作者:Qinghua Mao、Wanying Qu、Dadi Guo
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考