遇见数据集

junaid008/Pashto-100k-Pairs

收藏
Hugging Face2026-05-14 更新2026-05-31 收录
官方服务:

资源简介:

Qehwa AI Pashto 100K Fine-Tuning Dataset 是一个大规模普什图语指令调优数据集,包含超过10万条高质量指令-响应对,专为监督微调、对话人工智能和下游自然语言处理任务设计。该数据集旨在推动普什图语(一种全球数百万人使用但资源严重不足的语言)的人工智能研究,覆盖20多个多样化领域,并基于一个超过15亿普什图语词元的语料库进行整理。其主要目标是提升普什图语的语言理解、指令遵循、对话能力、推理与知识生成,以及大语言模型的领域适应能力。数据集采用JSONL格式,适用于指令调优、对话系统开发、NLP研究等用途。

Qehwa AI presents a large-scale Pashto instruction tuning dataset containing 100,000+ high-quality instruction-response pairs designed for supervised fine-tuning, conversational AI, and downstream NLP tasks. This dataset was created to advance AI research for the Pashto language, a significantly underrepresented low-resource language spoken by millions worldwide. It covers more than 20 diverse domains and is curated as part of a broader 1.5 billion token Pashto corpus. The primary goal is to improve Pashto language understanding, instruction following, conversational capabilities, reasoning and knowledge generation, and domain adaptation for large language models.

提供机构:
junaid008
二维码
社区交流群
二维码
科研交流群
商业服务