遇见数据集

African News Corpus

收藏
Zenodo2022-08-17 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

This consist of a monolingual news corpus for 19 languages from various sources like VOA, BBC, isolezwe etc. - The <strong>BBC</strong> corpus (except for Yoruba) was extracted from the AfriBERTa corpus, please cite the AfriBERTa paper if you use it. A big thank you to Kelechi Ogueji for providing this corpus - The <strong>VOA</strong> corpus was extracted from the MOT corpus, please cite the MOT paper if you use it. - The <strong>Isolezwe</strong> (<strong>xho</strong>, <strong>zul</strong>) was crawled as part of the Lacuna NER/POS project with Masakhane, please cite the MAFAND paper for that. - The <strong>nya</strong> data was part of the AI4D paper. - We thank Jonathan Mukiibi for providing the <strong>lug</strong> news corpus. - If you use the corpus for <strong>amh</strong>, <strong>hau</strong>, <strong>ibo</strong>, <strong>kin</strong>, <strong>lug</strong>, <strong>luo</strong>, <strong>pcm</strong>, <strong>swa</strong>, <strong>wol</strong>, <strong>yor</strong>, please cite our MAFT paper. We provide a description of the sources in the paper.

提供机构:
Zenodo
创建时间:
2022-08-14
二维码
社区交流群
二维码
科研交流群
商业服务