ai.hackcv
论文精选 82arXiv

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages· 多代理消息中的错误轨迹价值

Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likely to be correct is also worth keeping. Yet a wrong answer can contain a useful decomposition, constraint, or scientific principle. We test this distinction with Diverse Hypothesis Deliberation (DHD), a controlled measurement protocol that caches five independently generated messages and replays the same downstream solver, called the integrator, with each message available or hidden. The replay comparison measures a message's trajectory value: whether making the message available helps or harms subsequent reasoning. Across five mathematics and science benchmarks and two openly available model families, gpt-oss-120

领域:cs.AI作者:Chih-Hsuan Yang、Anjir Ahmed Chowdhury、Cheng-Hau Yang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考