ai.hackcv
论文精选 65arXiv

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting· 语言模型和深度研究代理的足球预测动态基准

Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models and deep-research agents. The 2026 FIFA World Cup is its first evaluation, and the same process can be reused for future leagues and cups. Before each match, a model either receives a common evidence package or searches for information itself. It predicts the result and score, likely players and events, match statistics, and the outcome of the competition. After the match, these predictions are compared with the recorded result. We report result accuracy, exact-score accuracy, and a scoreline score that gives some credit when a predicted score is

AI 解读论文

动态评估语言模型和代理在足球预测上的表现

核心方法
构建了一个动态基准WorldCupArena,使用2026年世界杯比赛前的实时信息作为输入,评估模型的预测准确性
适合谁读
研究者、工程师
要解决的问题
如何评估语言模型和代理在足球比赛预测中的能力,特别是在动态变化的信息环境中
关键实验
未提供
主要贡献
1. 提出了一个评估语言模型和代理在足球预测上表现的动态基准;2. 能够评估多个方面的预测,如比赛结果、比分、球员表现等
意义与局限
为语言模型和代理在体育赛事预测中的应用提供了重要评估工具,但仅限于足球领域,未来可以扩展到其他体育项目。模型的预测能力受到实时信息获取和处理能力的限制。
领域:cs.AI作者:Zhaokai Wang、Tianlin Gui、Jiayuan Rao
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考