team-dentaku/dentaku-stage1
收藏资源简介:
这是Team dentaku团队为FT-LLM2026数学任务调优竞赛构建提交系统时使用的数据集。该仓库发布了用于第一阶段监督微调(SFT)的数据集。数据集包含以下列(键):question(输入数学问题)、code(输出代码)和thinking(推理过程)。数据集通过使用openai/gpt-oss-120b从多种数据源合成,并经过多次基于规则和LLM的数据过滤和去重处理。最终数据集由6,094,013个实例组成。
This is the dataset used by Team dentaku to build their submission system for the FT-LLM2026 tuning competition (mathematics task). This repository releases the dataset used for the first-stage SFT dataset. The column names (keys) and their contents are as follows: question (the input mathematics problem), code (the output code), and thinking (the reasoning process). The dataset was synthesized from various data sources using openai/gpt-oss-120b, with rule-based and LLM-based data filtering and deduplication performed multiple times. The final dataset consists of 6,094,013 instances.



