Yankari
收藏资源简介:
Yankari数据集是由非洲语言保护中心创建的一个大规模单语种约鲁巴语数据集,旨在填补自然语言处理(NLP)领域中约鲁巴语资源的重大空白。该数据集包含51,407份文档,总计超过3000万Tokens,来源于13个不同的高质量来源,如新闻网站、博客和维基百科等。数据集的创建过程强调了伦理数据收集、严格的质量控制和语言真实性的保护,避免了宗教文本和机器翻译内容的使用。Yankari数据集的应用领域广泛,包括开发更精确的NLP模型、支持比较语言学研究以及促进约鲁巴语的数字可访问性。
The Yankari Dataset is a large-scale monolingual Yoruba dataset created by the African Language Conservancy, aimed at addressing the significant gap in Yoruba language resources within the field of natural language processing (NLP). It contains 51,407 documents totaling over 30 million Tokens, sourced from 13 distinct high-quality sources including news websites, blogs, Wikipedia, and others. The development of this dataset emphasizes ethical data collection, rigorous quality control, and the preservation of linguistic authenticity, excluding religious texts and machine-translated content. The Yankari Dataset has a wide range of applications, including developing more accurate NLP models, supporting comparative linguistics research, and promoting digital accessibility of the Yoruba language.

- 1Yankari: A Monolingual Yoruba Dataset非洲语言保护中心 · 2024年



