ai.hackcv
论文精选 65arXiv

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents· 超越成功率:成本感知的安全代理评估

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, we compare models at fixed cost levels and decompose performance by inference spend and tool spend. Our results show distinct scalingregimes for red- and blue-team tasks. Offensive CTF performance improves with additional test-time compute, and scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigation does not scale in the same way: success depends more heavily on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget alone. We argue that security-agent benchmarks should measure economic efficiency and operational fit alongside task success. Cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve. We present an interactive website with our results https://evals.frontier.security.

AI 解读论文

通过成本感知评估安全代理的进攻与防御性能

核心方法
在固定成本水平下,通过进攻性 Cybench 挑战和防御性 Splunk BOTS v1 调查挑战来评估语言模型安全代理的性能,并按推理成本和工具使用成本分解结果
适合谁读
研究者 / 工程师 / 产品经理
要解决的问题
现有的安全代理评估方法忽略了实际操作中的成本消耗,仅关注最佳情况下的成功率
关键实验
在进攻性 CTF 挑战和防御性 SOC 调查中进行了模型性能测试
主要贡献
揭示了红蓝团队任务不同成本收益模式;提出了结合经济效率和操作适应性的评估标准
意义与局限
对于理解和指导安全代理的实际应用具有重要价值,但也强调了防御代理在工具使用和数据选择上的局限性
领域:cs.CR作者:Paul Kassianik、Blaine Nelson、Yaron Singer
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考