ai.hackcv
论文精选 60arXiv

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change. To address this gap, we introduce PACE-Bench (Physics Adaptation via Code Evolution), a simulator-grounded benchmark of 144 source-to-target adaptation pairs across six physics domains. Each pair links a source environment to a mutated target environment with the same goal and interface. A code-driven design that succeeds in the source fails in the target, where agents must iteratively adapt it into a working target design using diagnostic sandbox feedback within a limited attempt budget. We compare ten self-evolving methods from four paradigms. The benchmark remains far from saturated:

AI 解读论文

针对动态环境中的物理适应性,提出了一种基于代码演化的基准测试。

核心方法
提出 PACE-Bench,一个包括144个源-目标适应对的基准测试,涵盖六个物理领域,通过有限次尝试和诊断性沙盒反馈来评估代理在目标环境中的适应能力。
适合谁读
研究者
要解决的问题
现有自我演进代理的评估方法通常在固定条件下优化,并未测试条件变化后的恢复能力。
关键实验
比较了来自四种范式的十种自我演进方法的性能,实验结果表明基准测试尚未达到饱和状态。
主要贡献
首次系统地评估了自我演进代理在动态环境中的物理适应性,为后续研究提供了基准测试平台。
意义与局限
意义在于推动自我适应代理的研究和发展,影响在于提供了一个新的评估标准,局限性在于目前仅有六种物理领域的适应对。
领域:cs.AI作者:Yuhao Zhan、Bingxiang He、Zecong Tang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考