遇见数据集

haoranli-ml/genvf-filtered-proof-graded_score7_only_with_summaries

收藏
Hugging Face2026-04-10 更新2026-04-12 收录
官方服务:

资源简介:

--- dataset_info: features: - name: index dtype: int64 - name: row_id dtype: int64 - name: problem dtype: string - name: answer dtype: 'null' - name: source list: string - name: mean_reward dtype: float64 - name: full_response dtype: string - name: full_reasoning dtype: string - name: model dtype: string - name: prefix dtype: string - name: prefix_end_index dtype: int64 - name: num_thoughts dtype: int64 - name: prefix_type dtype: string - name: prefix_type_description dtype: string - name: suffix_num list: int64 - name: suffix_model list: string - name: pending list: bool - name: pending_model list: 'null' - name: suffix_response list: string - name: suffix_summary list: string - name: self_summary list: string - name: suffix_reasoning list: string - name: finish_reason list: string - name: budget_used list: int64 - name: escalation list: int64 - name: usage list: - name: completion_tokens dtype: int64 - name: prompt_tokens dtype: int64 - name: total_tokens dtype: int64 - name: error list: 'null' - name: error_type list: 'null' - name: prefix_model dtype: string - name: gemini_summary_of_future dtype: string - name: gemini_summary_list list: string - name: prefix_steps list: string - name: suffix_variants list: - name: detailed_steps list: string - name: high_level_steps list: string - name: id dtype: int64 - name: dedup_note dtype: string - name: cross_prefix_alignment_scores list: - name: avg_alignment dtype: float64 - name: individual_scores list: - name: compared_row_id dtype: int64 - name: compared_summary_id dtype: int64 - name: direction dtype: string - name: output_text dtype: string - name: problem_index dtype: int64 - name: reasoning dtype: string - name: score dtype: float64 - name: num_comparisons dtype: int64 - name: summary_id dtype: int64 - name: filtered_suffix list: - name: detailed_steps list: string - name: high_level_steps list: string - name: id dtype: int64 - name: rubrics dtype: string - name: prefix_summary_steps dtype: string - name: filtered_suffix_summary_steps list: string - name: input_to_VF dtype: string - name: proof_scores list: - name: points dtype: int64 - name: suffix_id dtype: int64 - name: proof_details list: - name: assessment dtype: string - name: errors dtype: string - name: suffix_id dtype: int64 - name: partial_suffix list: string - name: suffix_cluster list: - name: cluster_id dtype: int64 - name: reasoning dtype: string - name: representative_index dtype: int64 - name: strategy_description dtype: string - name: suffix_indices list: int64 - name: prefix_summary dtype: string - name: detailed_suffix_summary list: string - name: high_level_suffix_summary list: string - name: dense_suffix_summary list: string - name: critical_moves_indices_in_suffix list: string splits: - name: train num_bytes: 308593746 num_examples: 1077 - name: test num_bytes: 7630108 num_examples: 26 download_size: 253123411 dataset_size: 316223854 configs: - config_name: default data_files: - split: train path: data/train-* - split: test path: data/test-* ---

提供机构:
haoranli-ml
搜集汇总
数据集介绍
haoranli-ml/genvf-filtered-proof-graded_score7_only_with_summaries 数据集图片
构建方式
在数学推理领域,高质量的证明生成与评估数据集对于推动自动推理系统的发展至关重要。该数据集通过多阶段筛选与评分机制构建,首先收集来自多样化数学问题的原始解答,随后利用语言模型生成多个推理路径,并引入验证机制对每个推理步骤进行严谨的评分。构建过程中特别注重证明的逻辑完整性与质量,仅保留评分达到特定阈值(如score7)的样本,同时为每个证明生成了结构化的摘要,确保了数据的高可靠性与逻辑严密性。
特点
该数据集的核心特征在于其精细的结构化标注与多层次的质量控制。每个样本不仅包含原始问题与生成的解答,还详细记录了推理过程、模型来源、奖励评分以及跨前缀对齐分数等元信息。数据集特别提供了证明的详细步骤与高级摘要,并集成了聚类分析与关键步骤索引,使得研究者能够深入分析不同推理策略的差异与有效性。这种多维度的特征设计为数学推理模型的训练与评估提供了丰富的监督信号。
使用方法
该数据集适用于数学自动推理与证明生成领域的研究,尤其适合用于训练或评估能够进行复杂逻辑推理的语言模型。使用者可通过加载指定的数据分割(如训练集或测试集)访问结构化字段,利用其中的问题、证明步骤、评分及摘要信息进行模型训练。此外,数据集提供的对齐分数与聚类信息可用于分析模型输出的多样性或进行对抗性验证,而摘要字段则能辅助实现更高效的推理过程监督或知识蒸馏。
背景与挑战
背景概述
在人工智能推理能力评估领域,生成多样化且高质量的推理轨迹对于推动模型逻辑思维与问题解决能力的发展至关重要。数据集“genvf-filtered-proof-graded_score7_only_with_summaries”应运而生,其构建旨在系统性地收集与标注由大型语言模型生成的数学证明或复杂问题求解过程,并通过严格的评分机制筛选出高质量样本。该数据集由研究团队精心设计,聚焦于探索模型在生成式验证框架下的推理性能,核心研究问题涉及如何有效评估与提升模型在结构化推理任务中的准确性与鲁棒性。其出现为自动化推理、教育技术及人工智能辅助证明等交叉领域提供了宝贵的基准资源,促进了可解释人工智能的进步。
当前挑战
该数据集致力于应对生成式验证中推理质量评估的挑战,核心在于确保模型生成推理轨迹的逻辑严谨性与正确性。具体挑战包括:在领域问题层面,如何设计普适且细粒度的评分标准以量化推理过程的优劣,以及如何处理数学证明中常见的抽象思维与步骤跳跃性;在构建过程中,需克服大规模生成数据的噪声过滤、多模型输出对齐的一致性校验,以及人工标注与自动评分相结合的可靠性与可扩展性难题。这些挑战共同指向了提升人工智能系统深层推理能力的核心瓶颈。
常用场景
经典使用场景
在人工智能推理与数学证明领域,该数据集通过整合问题描述、多模型推理路径及评分机制,为研究复杂逻辑推理过程提供了结构化资源。其经典使用场景聚焦于评估和比较不同大型语言模型在数学问题求解中的表现,特别是针对证明步骤的生成与验证。研究者可借助数据集中的详细推理链、奖励分数和总结信息,深入分析模型在生成严谨证明时的策略差异与能力边界,从而推动自动化推理技术的发展。
实际应用
在实际应用中,该数据集可服务于教育技术领域,辅助开发智能辅导系统,为学生提供个性化的数学证明指导与反馈。同时,在软件工程中,它能用于增强代码验证或定理证明工具的推理模块,提升自动化系统的逻辑严谨性。此外,数据集支撑的模型评估框架也为企业级AI解决方案的可靠性测试提供了基准,确保在金融、安全等高风险领域部署的推理模型具备足够的准确性与可解释性。
衍生相关工作
围绕该数据集,已衍生出一系列经典研究工作,主要集中在推理对齐、多模型协作证明生成以及自动化评分算法的开发上。例如,基于其结构化推理链的研究促进了思维链提示工程的优化;利用奖励分数进行的强化学习训练提升了模型的证明生成质量;而对推理步骤的聚类与总结分析则启发了新的可解释性方法。这些工作共同推动了人工智能在形式推理与逻辑验证领域的理论突破与应用扩展。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务