LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports· LLM-SoccerArena: 用 LLM 预测体育赛事的新基准
Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena (https://llm-soccerarena.com), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known. LLM-SoccerArena provides (1) a prospective live benchmark protocol, (2) a public open-source platform, and (3) a factorial benchmark design together with tournament-related questions (e.g., which team will win). LLM-SoccerArena automatically records timestamped,
LLM-SoccerArena:实时评估 LLM 在体育赛事预测中的表现
- 核心方法
- 构建了一个前瞻性的实时基准平台 LLM-SoccerArena,通过公开问题集(如比赛结果预测)来评估 LLM 的预测能力
- 适合谁读
- 研究者、工程师、体育分析师
- 要解决的问题
- 评估 LLM 在预测不确定的未来事件(如体育赛事)方面的能力
- 关键实验
- 未提供具体实验数据,但描述了平台的测试方法和流程
- 主要贡献
- 提供了实时基准协议、开源平台和因子基准设计,能够动态评估 LLM 的预测性能
- 意义与局限
- 为 LLM 在动态和不确定环境中的预测能力提供了一个新的评估工具,有助于推动模型的改进和应用;局限在于目前仅限于足球赛事