Dolci-Instruct-SFT-FR-Qwen3-30B-A3B
收藏资源简介:
Dolci Instruct SFT — FR (Qwen3-30B-A3B) 是一个经过领域标注的法语监督式微调(SFT)混合数据集。它源自 `allenai/Dolci-Instruct-SFT` 数据集,并利用 Qwen3-30B-A3B 模型进行了知识蒸馏。该数据集当前版本的主要更新在于:为每个样本新增了一个 `domain` 字段,该字段采用 AllenAI 的规范分类法,并根据样本的 `id` 信息重建而来;同时移除了原数据集中的“多语言”(包含 Aya 数据集,98,659 条)和“硬编码数据”(65 条)两个子集。数据集总共包含 1,807,651 个样本,涵盖 8 个不同的领域,具体分布为:编程(326,357 条)、对话(306,949 条)、推理(303,881 条)、数学(269,338 条)、其他(253,756 条)、精确指令遵循(134,003 条)、安全(110,126 条)和科学(103,241 条)。需要注意的是,父数据集中的“工具使用”子集未包含在此法语版本中。每个数据样本包含四个字段:唯一标识符 `id`、由 `content` 和 `role` 组成的对话消息列表 `messages`、整型的质量评分 `quality_score` 以及字符串类型的领域标签 `domain`。
Dolci Instruct SFT — FR (Qwen3-30B-A3B) is a domain-annotated French supervised fine-tuning (SFT) hybrid dataset. It is derived from the `allenai/Dolci-Instruct-SFT` dataset and underwent knowledge distillation using the Qwen3-30B-A3B model. The key updates in the current version of this dataset are as follows: a new `domain` field has been added to each sample, which adheres to AllenAI's standardized taxonomy and is reconstructed based on the sample's `id` information; additionally, two subsets from the original dataset, namely the "multilingual" subset (containing the Aya dataset with 98,659 entries) and the "hardcoded data" subset (consisting of 65 entries), have been removed. In total, this dataset contains 1,807,651 samples spanning 8 distinct domains, with the specific distribution as follows: Programming (326,357 entries), Conversation (306,949 entries), Reasoning (303,881 entries), Mathematics (269,338 entries), Other (253,756 entries), Precise Instruction Following (134,003 entries), Safety (110,126 entries), and Science (103,241 entries). It should be noted that the "Tool Use" subset from the parent dataset is not included in this French version. Each data sample contains four fields: a unique identifier `id`, a list of conversation messages `messages` composed of `content` and `role`, an integer-type quality score `quality_score`, and a string-type domain label `domain`.




