nassimjp/Pashto_OrcaCoT
收藏资源简介:
Pashto_OrcaCoT是一个高质量、开源的链式思维(CoT)数据集,专门为普什图语推理任务优化。该数据集基于OpenOrca数据集的普什图语翻译,包含多步逻辑和结构化推理模式,旨在弥补国际大型语言模型(LLM)基准与普什图语原生人工智能之间的差距。它作为iPashto.ai引擎的基础推理层,并为高级模型(如即将推出的Qwen3-1.7B-Pashto-Gold管道)提供关键对齐资源。数据集具有结构化链式思维、普什图语原生对齐(尊重语言细微差别和文化背景)、高质量完整性(减少幻觉、确保逻辑一致性)以及零爬取策略(通过精确策展提供生产就绪的指令调优)等特点。每个记录包含三个字段:instruction(系统提示或代理角色定义)、input(普什图语用户查询或任务陈述)和output(包含逐步推理逻辑和最终答案的标准响应)。该数据集旨在支持监督微调(SFT),提升普什图语语言模型在复杂推理任务上的性能,使其达到与英语等其他资源丰富语言相当的水平。
Pashto_OrcaCoT is the premier open-source, high-quality Chain-of-Thought (CoT) dataset natively optimized for the Pashto language. Built upon rigorous multi-step logic and structured reasoning patterns inspired by the Orca framework, this dataset bridges the gap between international LLM benchmarks and Pashto-native artificial intelligence alignment. It serves as the foundational reasoning layer for the iPashto.ai engine and acts as a critical alignment asset for advanced models, including the upcoming Qwen3-1.7B-Pashto-Gold pipeline. The dataset features structured CoT with multi-step logic flows, Pashto-native alignment that respects linguistic nuances and cultural subtexts, high-quality integrity to counteract hallucinations and ensure logical consistency, and a zero-scraping strategy for gold-standard instruction tuning. Each record follows a three-field schema: instruction (system prompt or agent persona), input (user query or task in Pashto), and output (gold-standard response with step-by-step reasoning and final answer). It is designed to enable Pashto language models to perform complex chain-of-thought deductions on par with English and other globally rich resource environments through supervised fine-tuning (SFT).




