AI-for-Education/Luganda-Linguistic-Pedagogical-Knowledge-Benchmark
收藏资源简介:
Luganda Linguistic Pedagogical Knowledge (LLPK) Benchmark是一个多项选择题基准数据集,用于衡量模型是否理解如何在乌干达背景下教授基础识字能力。该数据集通过混合LLM生成和人工审查流程构建,基于全球阅读科学(包括GEEAP报告)、教育部P1卢干达语教师指南和Pilkington的1915年《卢干达语手册》。数据集使用Gemini 2.5 Flash生成,并由教育专家和卢干达语母语专家验证。数据集包含两种语言配置:英语和卢干达语,每种语言有100个多项选择题。数据集包含多个列,如id(稳定的问题标识符)、section(教学主题)、skill(目标技能)、prompt(完整的多项选择题文本包括选项)、correct_answer(正确答案)、source_ref(语言/教学来源参考)和rationale(正确答案的解释)。数据集由AI for Education和Crane AI Labs合作构建。
The Luganda Linguistic Pedagogical Knowledge (LLPK) Benchmark is a multiple-choice benchmark measuring whether a model understands how to teach foundational literacy in the Ugandan context. Built via a hybrid LLM-generation + human-review pipeline grounded in the global reading science (incl. the GEEAP report), the Ministry of Education P1 Luganda Teachers Guide, and Pilkingtons 1915 *A Handbook of Luganda*. Generated with Gemini 2.5 Flash and validated by an education expert and native Luganda experts. 100 MCQs per language. The dataset includes columns such as id (stable question identifier), section (pedagogical theme), skill (specific skill targeted), prompt (full multiple-choice question text including options), correct_answer (single-letter correct answer), source_ref (references to linguistic / pedagogical sources), and rationale (why the correct answer is correct). Built by AI for Education as part of the Small Language Model finetuning project with Crane AI Labs.




