ai.hackcv
论文精选 60arXiv

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat· DungeonBench:Dungeons & Dragons 战斗中的战术推理基准

Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once. We introduce DungeonBench, a benchmark for tactical reasoning in Dungeons & Dragons combat, built to cover the vast majority of combat-relevant 2014 System Reference Document content whose effects can be resolved by the simulator while retaining mechanics that simplified combat simulators often abstract away. At each step, DungeonBench exposes a complete tactical observation, a pending decision, and an indexed list of executable options spanning movement, attacks, spells, reactions, objectives, preparation, and scarc

AI 解读论文

DungeonBench 为 Dungeons & Dragons 战斗设计战术推理基准。

核心方法
构建了 DungeonBench,一个涵盖了 2014 年系统参考文档中大部分战斗相关内容的基准,保留了简化战斗模拟器常常忽略的机制,为每一步决策提供完整的战术观察、待定决策和可执行选项的索引列表。
适合谁读
研究者、工程师
要解决的问题
当前的游戏和模拟器测试套件未能充分测试规则丰富的战术推理能力,尤其在多因素同时影响决策时的复杂情况。
关键实验
未提供
主要贡献
提供了首个全面测试 Dungeons & Dragons 战斗中战术推理能力的基准,能够评估代理在复杂规则环境下的决策能力。
意义与局限
DungeonBench 为研究复杂规则环境下的战术决策提供了新的工具,影响了 AI 在游戏和模拟器中的应用。局限在于仅局限于 Dungeons & Dragons 战斗场景,可能不适用于其他类型的战术推理。
领域:cs.AI作者:Ismayil Ismayilov、Atakan Kara、Kaan Oktay
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考