YembaEGRA
收藏资源简介:
The Yemba language is a Bantu language spoken in the western region of Cameroon. It is one of the ten languages spoken by the Bamileke peoples. The basic education system in Cameroon is made up of three levels: level 1: SIL-CEP, level 2: CE1-CE2, level 3: CM1-CM2 and lessons in “National Languages and Cultures” (introduced in 2019) are guided by a curriculum which for each level contains all the teaching units as well as eight thematic fields (or centers of interest) around which learning takes place. This corpus was built for learning automatic speech recognition models that can be used to facilitate the learning and assessment of national languages in the basic education system in Cameroon. The corpus of words available in this directory was formed for each center of interest by an educational facilitator who proposed a set of words. A linguist specializing in the Yemba language translated them to obtain a corpus of 60 words. These words were then pronounced twice by 69 native speakers, level 3 students including 36 girls and 33 boys. The recordings were carried out in classrooms and quiet rooms close to the public schools of Melah and Toudjoua (located in the village of Bamendou in the Menoua department, West region, Cameroon). In the metadata folder the corpus of words is present in a csv file named words_corpus. Information about each speaker is grouped in the speakers_description file in csv format including gender, age, class. The audio folder is divided into eight sub-folders named CI1 to CI8 each corresponding to a center of interest, within these folders we have three sub-folders named 1 to 3 for each level. Each of these subfolders contains the audio files of the words belonging to the center of interest and the level considered; These audios are grouped in subfolders named W1 to Wx (where x is the number of words of the center of interest). Each word folder contains audio files in wav format. Each audio file was named as follows: spkr__word__ occ__ci__l_.wav. For example, the files spkr_2_word_40_occ_1_ci_5_l_3.wav and spkr_2_word_40_occ_2_ci_5_l_3.wav correspond respectively to the files of occurrences 1 and 2 of word 40 belonging to center of interest 5, pronounced by speaker number 2 of level 3.
耶姆巴语(Yemba language)是一种通行于喀麦隆西部地区的班图语(Bantu language),为巴米莱克族群所使用的十种语言之一。喀麦隆基础教育体系分为三个阶段:第一阶段为SIL-CEP,第二阶段为CE1-CE2,第三阶段为CM1-CM2;2019年开设的“国家语言与文化”(National Languages and Cultures)课程,其教学大纲针对每个阶段涵盖全部教学单元,以及围绕学习活动展开的八大主题领域(或称兴趣中心)。本语料库旨在构建可用于助力喀麦隆基础教育体系中国家语言学习与测评的自动语音识别(automatic speech recognition)模型。本目录下的词语语料库,由教育辅导人员针对每个兴趣中心拟定词语集合,再由一名精通耶姆巴语的语言学家进行翻译,最终得到包含60个词语的语料库。随后,69名母语使用者(均为第三阶段学生,其中女生36名、男生33名)将每个词语朗读两遍。录音工作在靠近喀麦隆西部大区梅努阿省巴门杜村的梅拉(Melah)与图朱阿(Toudjoua)公立学校的教室及安静房间内完成。元数据文件夹中,词语语料库存储于名为words_corpus的CSV文件内。每位说话者的信息汇总于speakers_description的CSV文件,内容包含性别、年龄与所在年级。音频文件夹分为八个子文件夹,命名为CI1至CI8,分别对应一个兴趣中心;每个子文件夹下又设有1至3三个子文件夹,分别对应三个学习阶段。各阶段子文件夹内,按词语分组为W1至Wx(x为该兴趣中心下的词语总数),存放对应兴趣中心与学习阶段的词语音频文件。每个词语文件夹内均为WAV格式的音频文件。音频文件命名规则如下:spkr__word__occ__ci__l_.wav。示例:spkr_2_word_40_occ_1_ci_5_l_3.wav与spkr_2_word_40_occ_2_ci_5_l_3.wav,分别对应第三阶段2号说话者朗读的、属于兴趣中心5的第40个词语的第1、2次录音音频。



