perfect-pashto-reasoning-sft
收藏资源简介:
Perfect Pashto Reasoning SFT Dataset 是一个专为普什图语大语言模型(LLM)和AI社区设计的高质量监督微调(SFT)数据集。该数据集基于英文的Magpie-Pro-300K-Filtered数据集构建,旨在提升普什图语模型(如Rawan和Ghanam系列)的深度推理与思维链能力。数据经过原子级、逐行的翻译、高质量过滤(包括MD5哈希去重)、智能分块重组以及严格的自然语言和逻辑流程检查,最终重新格式化为与Hugging Face Alignment Handbook完全兼容的结构。数据集共包含300,000个样本,划分为270,000个训练样本和30,000个测试样本。每个样本采用JSONL格式,包含三个核心字段:`system`(定义模型身份和推理框架)、`instruction`(普什图语用户指令或问题)以及`output`(高质量、流畅的普什图语逻辑回答)。该数据集适用于监督微调、思维链与深度推理、普什图语指令遵循、学术与技术写作以及合成数据生成等任务,并在Apache-2.0许可证下发布,允许商业、学术和研究用途。
Perfect Pashto Reasoning SFT Dataset is a high-quality supervised fine-tuning (SFT) dataset specifically designed for Pashto large language models (LLMs) and the AI community. This dataset is constructed based on the English Magpie-Pro-300K-Filtered dataset, aiming to enhance the deep reasoning and chain-of-thought capabilities of Pashto language models such as the Rawan and Ghanam series. The data undergoes atomic-level, line-by-line translation, high-quality filtering including MD5 hash deduplication, intelligent chunking and recombination, as well as strict natural language and logical flow checks, before being finally reformatted into a structure fully compatible with the Hugging Face Alignment Handbook. The dataset contains a total of 300,000 samples, divided into 270,000 training samples and 30,000 test samples. Each sample is in JSONL format and includes three core fields: `system` (defines the model identity and reasoning framework), `instruction` (Pashto user instructions or questions), and `output` (high-quality, fluent Pashto logical responses). This dataset is suitable for tasks such as supervised fine-tuning, chain-of-thought and deep reasoning, Pashto instruction following, academic and technical writing, and synthetic data generation, and is released under the Apache-2.0 license, permitting commercial, academic and research uses.
数据集概述
Perfect Pashto Reasoning SFT Dataset 是一个专为普什图语(Pashto)大语言模型设计的高质量监督微调(SFT)与深度推理(Deep Reasoning)数据集。
- 许可证: Apache-2.0
- 任务类别: 文本生成、问答
- 语言: 普什图语 (ps)
- 数据规模: 300,000 条样本(100K < n < 1M)
- 数据集划分:
- 训练集 (train_sft): 270,000 条
- 测试集 (test_sft): 30,000 条
- 数据格式: JSONL(兼容 Hugging Face Alignment Handbook)
数据集构建过程
数据集基于 Magpie-Pro-300K-Filtered 构建,经过以下步骤处理:
- 源头选择: 选用高质量的英文合成推理数据。
- 原子翻译: 对每对指令和输出进行逐行翻译,确保对齐无误。
- 高质量过滤: 通过 MD5 哈希去除重复和低质量样本。
- 分块与重组: 将长回答逻辑分段后,重新组合为流畅文本。
- 清洗与丰富: 人工和自动检查普什图语的自然性、语法正确性和逻辑流程。
- 格式重组: 将数据转换为适配 Alignment Handbook 的结构(
system,instruction,output字段)。 - 训练/测试分割: 按 90%(27万条)和 10%(3万条)划分。
数据记录结构
每条记录包含三个字段:
system: 模型身份和推理框架描述instruction: 用户的普什图语提问或任务output: 标准、流畅、高质量的普什图语逻辑回答
示例记录
json { "system": "تاسو یو هوښیار، ریښتینی او مرستندوی پښتو ژبی د ارادې او عمیق استدلال (Reasoning) ماډل یاست.", "instruction": "د ستورپوهنې د ساینس په اړه ځینې په زړه پورې حقایق څه دي؟", "output": "ستورپوهنه یوه په زړه پورې او پراسرار ساینسي برخه ده! ..." }
应用领域
- 监督微调 (SFT)
- 思维链与深度推理 (Chain-of-Thought & Deep Reasoning)
- 普什图语指令跟随
- 学术、技术与创意写作
- 合成数据生成





