When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit· 共享回放失败的防御性驾驶评估审计
Defensive driving scores are useful only when they preserve distinctions between policies that observe surrounding actors and those that do not. Re-simulation benchmarks may use reference-conditioned forgiveness, under which an agent receives credit when the logged human reference fails a compliance channel. When agent and reference share an unstable rollout transformation, this rule can propagate shared reference failures into broad compliance credit. We audit this risk in NAVSIM v2.2 original scene single-stage scoring. Under the affected documented-stack condition on the audited numerical backend, the route-blind Ignore-All probe and a route-aware actor-blind probe outrank human replay and PDM-Closed over the complete 12,146-token navtest split. A fresh installation following the public
审计 NAVSIM v2.2 中共享回放在防御性驾驶评分中的失败风险。
- 核心方法
- 通过在 NAVSIM v2.2 中对原始场景单阶段评分进行审计,分析不同策略下的评分表现,特别是忽略所有信息和忽略演员信息的策略与人类回放和 PDM-Closed 的对比。
- 适合谁读
- 研究者 / 工程师
- 要解决的问题
- 论文讨论了在防御性驾驶评估中,当代理和参考共享不稳定回放转换时,参考条件下的宽恕规则可能导致广泛的合规评分问题。
- 关键实验
- 在 12,146 个导航测试样例上,使用不同的策略(包括 Ignore-All 和 actor-blind)与人类回放和 PDM-Closed 进行对比测试。
- 主要贡献
- 揭示了参考条件下的宽恕规则在防御性驾驶评估中的潜在缺陷,以及共享回放的不稳定性和评分体系的关系。
- 意义与局限
- 这一发现对如何设计更公平、更准确的自动驾驶车辆评估系统具有重要意义,但仅限于特定版本的 NAVSIM,可能不适用于其他系统。