FinanceReasoning
收藏资源简介:
FinanceReasoning是一个专为评估大型推理模型在金融数值推理问题中的推理能力而设计的基准数据集。该数据集由北京邮电大学的研究团队创建,包含2238个问题,覆盖了广泛的金融知识,旨在解决现有数据集在金融领域知识覆盖面、推理复杂性和准确性方面的不足。每个问题都包括混合上下文、明确的问题、Python格式的解决方案和精确答案,为准确评估大型推理模型的复杂数值推理能力提供了可靠的参考。此外,还构建了一个包含3133个Python格式函数的全面金融函数库,以增强模型的金融推理能力。
FinanceReasoning is a benchmark dataset specifically designed to evaluate the reasoning capabilities of large reasoning models on financial numerical reasoning problems. Developed by the research team from Beijing University of Posts and Telecommunications, this dataset contains 2,238 questions covering a wide range of financial knowledge, aiming to address the shortcomings of existing datasets in terms of financial domain knowledge coverage, reasoning complexity and accuracy. Each question includes mixed context, explicit problem statement, Python-formatted solution and precise answer, providing a reliable reference for accurately evaluating the complex numerical reasoning capabilities of large reasoning models. In addition, a comprehensive financial function library containing 3,133 Python-formatted functions has been constructed to enhance the financial reasoning abilities of the models.
FinanceReasoning 数据集概述
基本信息
- 数据集名称: FinanceReasoning
- 研究机构: 北京邮电大学
- 会议: ACL 2025 Main Conference
- 论文状态: ACL 2025主会议论文
- 代码与资源: 提供arXiv论文和代码链接
核心贡献
-
可信度提升:
- 更新了4个公开数据集中15.6%的问题
- 新增908道带Python详细解法的问题
- 建立了严格的评估标准
-
全面性增强:
- 覆盖67.8%的金融概念和公式
- 构建3,133个Python格式化函数
- 显著提升模型表现(如GPT-4o从83.2%→91.6%)
-
挑战性设计:
- 包含238道高难度问题
- 需要应用多重金融公式进行精确数值推理
- 当前最佳模型(OpenAI o1+PoT)准确率89.1%
关键技术指标
- 知识覆盖率: 计算问题涉及的金融计算占金融百科全书比例
- 重标注比例: 测试/验证集样本更新比例9.7%-30%
- 性能提升: DeepSeek-RT较DeepSeek-V3有显著改进
模型要求
- 需基于给定条件(如最低预期回报率)选择合适金融公式
- 执行带舍入要求的逐步精确数值计算




