African News Corpus
收藏资源简介:
This consist of a monolingual news corpus for 19 languages from various sources like VOA, BBC, isolezwe etc. - The <strong>BBC</strong> corpus (except for Yoruba) was extracted from the AfriBERTa corpus, please cite the AfriBERTa paper if you use it. A big thank you to Kelechi Ogueji for providing this corpus - The <strong>VOA</strong> corpus was extracted from the MOT corpus, please cite the MOT paper if you use it. - The <strong>Isolezwe</strong> (<strong>xho</strong>, <strong>zul</strong>) was crawled as part of the Lacuna NER/POS project with Masakhane, please cite the MAFAND paper for that. - The <strong>nya</strong> data was part of the AI4D paper. - We thank Jonathan Mukiibi for providing the <strong>lug</strong> news corpus. - If you use the corpus for <strong>amh</strong>, <strong>hau</strong>, <strong>ibo</strong>, <strong>kin</strong>, <strong>lug</strong>, <strong>luo</strong>, <strong>pcm</strong>, <strong>swa</strong>, <strong>wol</strong>, <strong>yor</strong>, please cite our MAFT paper. We provide a description of the sources in the paper.



