ai.hackcv
论文精选 60arXiv

GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks· GeoBenchLLM: 评估大模型地理任务

In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo-related tasks and domains, and evaluate a set of LLMs on geo-spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.

AI 解读论文

LLM地理任务评估新基准:GeoBenchLLM

核心方法
构建包含12个公开数据集的综合基准,涵盖多样地理任务,评估多个LLM的地理空间与时间理解能力
适合谁读
研究者、工程师
要解决的问题
评估现有大语言模型在地理相关任务上的泛化能力
关键实验
未提供具体实验数据,但评估了多个LLM在不同地理任务上的表现
主要贡献
提供了一个新的、全面的基准工具,展示了模型规模和推理能力对性能的显著影响
意义与局限
有助于深入了解LLM在地理任务中的应用潜力及局限,为模型优化提供方向
领域:cs.AI作者:Rodrigo Ferreira Rodrigues、Karim Radouane、Jose G Moreno
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考