InstructGpt-educational
收藏资源简介:
LuminaSFT 是一个专为小型语言模型(SLMs)设计的合成监督微调(SFT)数据集集合,通过教师引导的数据再生和任务特定的合成数据生成方法创建。该集合包含五个子数据集:UltraChat200K-regenerated(通用指令数据再生)、InstructGpt-NaturalQa(事实问答)、InstructGpt-TriviaQa(事实问答)、Cot-Drop(阅读理解)和InstructGpt-educational(教育问答)。其中,InstructGpt-educational 子数据集包含三个文件,完全通过结构化多步提示生成,未使用种子数据。所有数据均使用先进的教师模型(如 DeepSeek-V3 和 Qwen/Qwen3-30B-A3B-Instruct-2507)生成,适用于文本生成、问答等自然语言处理任务。数据集采用 Open RAIL-D 许可证发布。
LuminaSFT is a collection of synthetic supervised fine-tuning (SFT) datasets specifically designed for small language models (SLMs), created via teacher-guided data regeneration and task-specific synthetic data generation methods. This collection includes five sub-datasets: UltraChat200K-regenerated (general instruction data regeneration), InstructGpt-NaturalQa (fact-based question answering), InstructGpt-TriviaQa (fact-based question answering), Cot-Drop (reading comprehension), and InstructGpt-educational (educational question answering). Among them, the InstructGpt-educational sub-dataset contains three files, which are entirely generated via structured multi-step prompts without using seed data. All data is generated using state-of-the-art teacher models such as DeepSeek-V3 and Qwen/Qwen3-30B-A3B-Instruct-2507, and is applicable to natural language processing (NLP) tasks including text generation and question answering. The dataset is released under the Open RAIL-D license.




