ai.hackcv
论文精选 65arXiv

V-FiLLM: Verified Financial LLM Reasoning Benchmark· V-FiLLM: 金融大模型推理评测基准

While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored. We introduce V-FiLLM, a framework that generates financial reasoning benchmarks from executable computation trees grounded in real tables, yielding items whose answers are correct by construction. Trees are evaluated symbolically to obtain ground truth and rendered into natural-language questions, removing any model from the labeling loop, so items can be generated at arbitrary scale without annotation cost and without inheriting a generator's error rate. V-FiLLM exposes four independently controllable axes of difficulty including computation depth, expression breadth, financial concept complexity, and context size. B

AI 解读论文

提出 V-FiLLM,金融大数据模型推理评测框架,无需标注成本生成大规模准确测试题。

核心方法
通过执行计算树生成自然语言问题,确保答案正确,同时控制四个难度维度:计算深度、表达宽度、金融概念复杂度和上下文大小。
适合谁读
研究者、金融工程师
要解决的问题
现有评测基准在金融领域结构化数据上的推理能力评估不足。
关键实验
未提供
主要贡献
提供了首个基于实际表格生成的金融推理评测基准,支持大规模无标注成本的题目生成。
意义与局限
有助于评估和改进大模型在金融领域的推理能力,但目前缺乏实验数据验证。
领域:cs.AI作者:Alicia Larsen、Victoire Laurent、Aulia Kharis Rakhamsari
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考