EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards· EarthVerse:跨动态地球系统和自然灾害的科学智能体基准测试
Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natural hazards make this work consequential because incomplete evidence can change estimates of severity, exposure, and mechanism. We introduce EarthVerse, a benchmark that evaluates scientific agents through package-scoped investigations. Its 405 reproducible tasks are grounded in 199 documented events and 19 hazard families. Agents inspect heterogeneous event packages, choose compatible evidence, execute transparent calculations, reconcile source differences, and preserve provenance in the final answer. We provide executable ground truth that decomposes each task into fine-grained answer units, together with task-specific rubrics that assess the supporting
提出动态地球系统和自然灾害分析的基准测试平台EarthVerse。
- 核心方法
- 通过构建包括405个基于199个文档事件和19个灾害家族的可重现任务的基准测试,EarthVerse平台要求智能体检查异构事件包,选择合适证据,执行透明计算,协调来源差异并保留最终答案的来源。
- 适合谁读
- 研究者
- 要解决的问题
- 该论文旨在解决地球系统分析中由于数据源、规模、时间及模式差异导致的自然危害评估不准确问题。
- 关键实验
- 关键实验包括对各任务的执行与评估,以及智能体处理异构数据集的能力测试。
- 主要贡献
- 贡献了一个全面的基准测试平台,用于评估处理跨动态地球系统和自然灾害任务的科学智能体的性能。
- 意义与局限
- EarthVerse为地球科学智能体的研究与开发提供了一个重要工具,有助于提高自然灾害预测和响应的准确性。然而,该平台目前可能局限于特定类型的科学智能体,未来需要更多的智能体参与测试以全面评估其效能。