ai.hackcv
论文精选 65arXiv

How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation· 自动化事实核查系统的鲁棒性如何?

Automated fact-checking (AFC) systems retrieve evidence and predict claim veracity, yet evaluations omit simple baselines, systems are developed for a single benchmark and cannot be trusted to generalise across domains. No prior work cross-evaluates the full two-stage retrieve-then-verify pipeline across diverse datasets, complementing retrieval-only studies (Thakur et al., 2021) and single-stage benchmarking studies (Calamai et al., 2025). We benchmark nine models, ranging from random and sparse baselines to fine-tuned transformers, zero-shot LLMs, and the two highest-ranked systems from the AVeriTeC 2025 shared task, across four datasets spanning scientific, open-web, and climate domains. Three findings stand out: (1) on ClimateCheck claim-only and fine-tuned models outperform zero-shot

AI 解读论文

评估自动化事实核查系统的跨域鲁棒性。

核心方法
对从随机基线模型到微调变换器和零样本LLMs在内的九种模型,进行跨四个不同领域数据集的基准测试。
适合谁读
研究者、工程师
要解决的问题
自动化事实核查系统在不同领域内的泛化能力不足,缺乏跨域评估。
关键实验
关键实验包括对科学、开放网络和气候变化等四个领域数据集上的九个模型进行对比测试。
主要贡献
首次对自动化事实核查的两阶段流程进行全面跨域评估,揭示了不同模型在特定数据集上的表现差异。
意义与局限
研究有助于理解自动化事实核查系统的局限性和鲁棒性,为未来系统开发和改进提供重要参考。但测试数据集有限,结论可能不完全普适。
领域:cs.AI作者:Aida Usmanova、Zangir Iklassov、Markus Leippold
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考