ai.hackcv
论文精选 82arXiv

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents· 行动胜于言语:评估跨语言政策保留

When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how the system fails, and are the only auditable part of its behaviour. We make the action policy the measured object across 8 models, 6 parallel benchmarks and 41 languages (2.38M rollouts). The naive measurement fails: five confounds sit between raw trace similarity and any defensible claim, each able to flip a conclusion. Short traces score higher, empty traces score perfectly, unrelated traces agree by chance over half the time, the gap is capped by each model's reproducibility, and a model asked the same question twice in on

领域:cs.CL作者:Sourabrata Mukherjee、Kalika Bali、Sunayana Sitaram
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考