instruct-pashto-h-school
收藏资源简介:
Instruct Pashto School 是一个普什图语(Pashto)教育问答指令数据集,专门为训练普什图语指令跟随大语言模型而设计。该数据集源自NCERT风格的教育内容,包含学校6至12年级多个学科的标准化问答对。数据集采用JSONL格式,遵循Ministral-Instruct模式,每条记录包含用户提问和助手回答的简单对话结构,无额外元数据,且已进行去重和清理处理,可直接用于监督微调。数据集规模为112,604个唯一问答对,全面覆盖科学(6-10年级)、生物学、化学、物理学、经济学(11-12年级)、商业研究、会计学、地理、历史、政治科学、社会学以及社会科学等多个学科领域。其主要用途包括:开发普什图语教育大语言模型、学校科目辅导、通用知识问答系统、指令跟随模型训练以及多语言教育研究。数据集基于MIT许可证发布,原始内容来源于相关作者的宽松许可数据集,转换和结构化工作由Nsibullah Nassim完成。
Instruct Pashto School is a Pashto educational question-answering instruction dataset specifically designed for training Pashto instruction-following large language models. This dataset is derived from NCERT-style educational content and contains standardized question-answering pairs across multiple subjects for grades 6 to 12 in schools. The dataset is in JSONL format and follows the Ministral-Instruct schema. Each record features a simple dialogue structure consisting of a user query and an assistant response, with no additional metadata. It has been deduplicated and cleaned, and is ready for direct use in supervised fine-tuning. The dataset has a scale of 112,604 unique question-answering pairs, comprehensively covering multiple subject areas including Science (Grades 6-10), Biology, Chemistry, Physics, Economics (Grades 11-12), Business Studies, Accounting, Geography, History, Political Science, Sociology, and Social Sciences. Its main applications include: developing Pashto educational large language models, school subject tutoring, general knowledge question-answering systems, instruction-following model training, and multilingual educational research. The dataset is released under the MIT License. Its original content is sourced from a loosely licensed dataset by the relevant authors, and the conversion and structuring work was completed by Nsibullah Nassim.





