OpenLLM-France/Luciole-PostTraining-Dataset-1.1
收藏资源简介:
Luciole-PostTraining-Dataset-1.1是一个经过精心策划的开放指令式文本数据集合,专为语言模型后训练设计。该数据集包含混合的合成与非合成指令,用于监督微调(SFT)以及为偏好对齐(如DPO)设计的响应对。除安全对齐数据外,对齐对均通过增量学习方法合成生成:所有对均使用Qwen3-32B和Qwen3-0.6B生成,其中前者被标记为接受响应。安全数据通过混合使用Qwen3-14B、Ministral-3-14B-Instruct以及Luciole-8B-Instruct的中间检查点(在SFT后)生成。这些响应对经过Ministral-14B-Reasoning和Qwen3-14B双重判断,仅当两个模型在标签上一致时才被纳入。尽管包含一些多语言数据,但数据集主要使用英语。该数据集由OpenLLM France联盟在法国BPI资助的France 2030项目下创建,旨在促进符合开源要求和欧洲AI开发及知识产权法律的大型语言模型训练。数据集分为四个子集:sft_instruct(无思维痕迹的指令数据)、sft_thinking(带思维痕迹的指令数据)、dpo_instruct(无思维痕迹的接受与拒绝响应对)和dpo_thinking(带思维痕迹的响应对,即将推出)。
The Luciole-PostTraining-Dataset-1.1 is a curated collection of open, instruction-style text data designed for language model post training. It includes a mixture of synthetic and non-synthetic instructions for supervised fine-tuning (SFT) as well as pairs of responses designed for preference alignment (e.g., DPO). With the exception of data for safety alignment, the alignment pairs were generated synthetically with a delta learning approach: all pairs were generated with Qwen3-32B and Qwen3-0.6B and the former were labeled as the accepted responses. Safety data were generated with a mixture of Qwen3-14B, Ministral-3-14B-Instruct, and an interim checkpoint of Luciole-8B-Instruct, after SFT. Pairs were judged with both Ministral-14B-Reasoning and Qwen3-14B. A pair was included only if both models agreed on their labels. While it contains some multilingual data, it is primarily English. The dataset was created by the consortium of the OpenLLM France project funded by BPI France as part of the France 2030 program, aiming to facilitate training of large language models in conformance with open-source requirements and European laws on AI development and intellectual property. It is divided into four subsets: sft_instruct (instruction-style data without thinking traces), sft_thinking (instruction-style data with thinking traces), dpo_instruct (instructions with pairs of accepted and rejected responses without thinking traces), and dpo_thinking (instructions with pairs of accepted and rejected responses with thinking traces, coming soon).




