Beyond Component Testing: Validating Agentic AI Systems· 超越组件测试:评估代理型AI系统
Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches validation practice beyond component testing and one-shot input--output evaluation, because acceptable system behavior now depends on how decisions unfold over time and under changing environmental conditions. This survey synthesizes 257 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems. The review is organized around a five-dimension taxonomy covering behavioral, safety, temporal, regulatory, and multi-agent concerns, and uses that taxonomy to map current approaches and expose recurrent coverage gaps. The
评估代理型AI系统的动态行为与安全性
- 核心方法
- 提出一个五维分类法,涵盖行为、安全、时间、监管和多代理问题,以系统地评估代理型AI系统
- 适合谁读
- 研究者
- 要解决的问题
- 传统组件测试和一次性输入输出评估方法无法充分验证代理型AI系统的长期动态行为和适应性
- 关键实验
- 未提供
- 主要贡献
- 总结现有评估方法,识别覆盖不足的领域,并为未来研究提供方向
- 意义与局限
- 为评估复杂代理型AI系统提供了框架,有助于提高系统可靠性,但需更多实验验证