lamm-mit/scientific-sft-grpo-data
收藏资源简介:
该数据集包含多个配置,主要用于训练和评估自然语言处理模型,特别是针对任务完成、科学设计推理和指令遵循等场景。数据集特征包括任务ID、论文ID、提示(包含角色和内容)、任务描述、参考推理、参考答案、参考证据、机制步骤、因果链接、必需概念、必需约束、评估标准、可接受替代方案、失败模式等。数据集划分为训练集、验证集和测试集,不同配置的数据量从几千字节到几千万字节不等,示例数量从几个到几千个。数据集可能用于支持强化学习(如GRPO)和监督微调(SFT)等机器学习方法,适用于生成对话、科学问题解决和因果推理等任务。
This dataset includes multiple configurations designed for training and evaluating natural language processing models, particularly for task completion, scientific design reasoning, and instruction following. The features consist of task ID, paper ID, prompt (with role and content), task description, reference reasoning, reference answer, reference evidence, mechanism steps, causal links, required concepts, required constraints, evaluation criteria, acceptable alternatives, failure modes, and more. The dataset is split into training, validation, and test sets, with data sizes ranging from a few kilobytes to tens of megabytes and example counts from a few to thousands. It is likely used to support machine learning methods such as reinforcement learning (e.g., GRPO) and supervised fine-tuning (SFT), applicable to tasks like dialogue generation, scientific problem-solving, and causal reasoning.
数据集概览:lamm-mit/scientific-sft-grpo-data
- 数据集名称:lamm-mit/scientific-sft-grpo-data
- 发布机构:LAMM: MIT Laboratory for Atomistic and Molecular Mechanics (lamm-mit)
- 模态:文本 (Text)
- 格式:parquet, optimized-parquet
- 规模:10K - 100K 条记录
- 库依赖:Datasets, pandas, Polars 等
数据集内容
该数据集包含科学机制推理任务,每个样本包含以下字段:
- task_id: 任务唯一标识符
- paper_id: 来源论文ID
- prompt: 系统与用户提示消息,要求解决科学机制任务并解释因果链
- task: 科学机制任务描述文本
- reference_reasoning: 参考推理过程
- reference_answer: 参考答案
- reference_evidence: 任务中引用最相关观察的原文
- mechanism_steps: 机制步骤列表
- causal_links: 因果链列表(包含因果关系键值对)
- required_concepts: 所需概念列表
- source_license: 来源许可证(如 CCBY)
- source_url: 来源论文URL
数据集划分与子集
- 子集 (Subset): 共8个,包括
grpo(17行)、scientific_design_grpo(675行)、scientific_design_grpo_L(1.2k行)、scientific_design_sft(575行)、scientific_design_sft_L(10k行)、scientific_design_sft_L_messages(10k行)、scientific_design_sft_messages(575行)、sft(20行) - 划分 (Split): 训练集 (train) 8行、验证集 (validation) 4行、测试集 (test) 5行(示例中仅展示grpo子集的train划分)
示例数据内容
数据集样本涵盖跨学科科学机制,例如:
- 磁力计中微磁体与聚合物变形对光学腔响应的影响
- 急性缺血性中风(AIS)的院前筛查与治疗黄金时间
- 内毛细胞突触丢失与代偿性增强导致听觉过敏
- 一氧化氮合酶缺失后其他亚型补偿促进骨愈合
- 家庭血糖监测与低血糖报告之间的关联(横断面研究局限性)
- 卫星辐射测量不确定度传播至气候数据记录
- 治疗性大麻使用动机与未满足医疗需求
- microRNA319通过切割位点调控转录因子TCP基因表达




