CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations· 电信网络运维中AI代理的故障排查能力评估
Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints. However, existing evaluations fail to accurately model real network characteristics or assess agents under partially observable telecom environments with diverse vendors, devices, protocols, and interfaces. In this paper, we introduce CTBench, a public benchmark for assessing whether an agent behaves like a competent telecom troubleshooting engineer. CTBench focuses on root cause analysis and path restoration. Each task is constructed by experts and annotated with rich task metadata, including golden evidence steps. CTBench uses expert-grounded metr
CTBench评估AI代理在电信网络运维中的故障排查能力。
- 核心方法
- 引入CTBench,一个公开的基准测试平台,涵盖根因分析和路径恢复任务,由专家构建并标注丰富的任务元数据。
- 适合谁读
- 研究者、工程师
- 要解决的问题
- 现有评估方法无法准确模拟真实网络特性,也不能评估AI代理在部分可观测的电信环境中处理多种厂商、设备、协议和接口的能力。
- 关键实验
- 未提供
- 主要贡献
- CTBench提供了一个评估AI代理是否像合格的电信故障排查工程师行为的工具。
- 意义与局限
- CTBench有助于推动AI在电信网络运维中的实际应用,但目前缺乏实际实验数据来验证其有效性。