GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks· GeoBenchLLM: 评估大模型地理任务
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo-related tasks and domains, and evaluate a set of LLMs on geo-spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.
LLM地理任务评估新基准:GeoBenchLLM
- 核心方法
- 构建包含12个公开数据集的综合基准,涵盖多样地理任务,评估多个LLM的地理空间与时间理解能力
- 适合谁读
- 研究者、工程师
- 要解决的问题
- 评估现有大语言模型在地理相关任务上的泛化能力
- 关键实验
- 未提供具体实验数据,但评估了多个LLM在不同地理任务上的表现
- 主要贡献
- 提供了一个新的、全面的基准工具,展示了模型规模和推理能力对性能的显著影响
- 意义与局限
- 有助于深入了解LLM在地理任务中的应用潜力及局限,为模型优化提供方向