遇见数据集

AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54

收藏
Hugging Face2026-04-27 更新2026-05-03 收录
官方服务:

资源简介:

NOESIS DORA SFT数据集是一个多语言监督微调数据集,专为NOESIS QwQ+DeepSeek-R1混合专家(MoE)管道构建。该数据集包含100万条记录和一个5万条精选高质量子集,支持30多种语言(如英语、中文、俄语、阿拉伯语、印地语、西班牙语等),涵盖文本生成任务,包括推理、代码、数学、心理学等领域,并包含思维链(chain-of-thought)痕迹。数据格式为JSONL,每条记录包含用户问题和助手回答,其中部分记录带有“<think>...</think>”推理块。数据来源于多个AI模型生成或蒸馏的输出,如Claude Sonnet 4.6、Claude Opus 4.7、DeepSeek-R1、DeepSeek V4、Qwen3.6、Gemini 3.1和GPT-5.4。数据集旨在用于DoRA SFT微调、路由器微调(用于CMoE架构)以及多语言指令调优与推理痕迹蒸馏,主要目标为NOESIS-QwQ-R1管道(QwQ-32B + DeepSeek-R1-32B TIES合并→CMoE 16E)。数据集遵循Apache 2.0许可证,部分数据基于Aya数据集(Apache 2.0)和NOESIS合成数据。

The NOESIS DORA SFT Dataset is a multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 Mixture of Experts (MoE) pipeline. It contains 1 million records and a curated subset of 50,000 high-quality records, supporting over 30 languages (e.g., English, Chinese, Russian, Arabic, Hindi, Spanish, etc.) and covering text-generation tasks including reasoning, code, math, psychology, and more, with chain-of-thought traces. The data format is JSONL, where each record includes a user question and assistant answer, some containing <think>...</think> reasoning blocks. The dataset is sourced from outputs generated or distilled by multiple AI models such as Claude Sonnet 4.6, Claude Opus 4.7, DeepSeek-R1, DeepSeek V4, Qwen3.6, Gemini 3.1, and GPT-5.4. It is intended for DoRA SFT fine-tuning, router fine-tuning (for CMoE architectures), and multilingual instruction tuning with reasoning trace distillation, primarily targeting the NOESIS-QwQ-R1 pipeline (QwQ-32B + DeepSeek-R1-32B TIES merge → CMoE 16E). The dataset is licensed under Apache 2.0, with portions derived from the Aya dataset (Apache 2.0) and NOESIS synthetic data.

提供机构:
AMAImedia
二维码
社区交流群
二维码
科研交流群
商业服务