junaid008/Pashto-100k-Pairs
收藏资源简介:
Qehwa AI Pashto 100K Fine-Tuning Dataset 是一个大规模普什图语指令调优数据集,包含超过10万条高质量指令-响应对,专为监督微调、对话人工智能和下游自然语言处理任务设计。该数据集旨在推动普什图语(一种全球数百万人使用但资源严重不足的语言)的人工智能研究,覆盖20多个多样化领域,并基于一个超过15亿普什图语词元的语料库进行整理。其主要目标是提升普什图语的语言理解、指令遵循、对话能力、推理与知识生成,以及大语言模型的领域适应能力。数据集采用JSONL格式,适用于指令调优、对话系统开发、NLP研究等用途。
Qehwa AI presents a large-scale Pashto instruction tuning dataset containing 100,000+ high-quality instruction-response pairs designed for supervised fine-tuning, conversational AI, and downstream NLP tasks. This dataset was created to advance AI research for the Pashto language, a significantly underrepresented low-resource language spoken by millions worldwide. It covers more than 20 diverse domains and is curated as part of a broader 1.5 billion token Pashto corpus. The primary goal is to improve Pashto language understanding, instruction following, conversational capabilities, reasoning and knowledge generation, and domain adaptation for large language models.



