Apakose-Ezekiel/farahan-yoruba-elder-corpus
收藏资源简介:
Farahàn Yoruba Elder Corpus 是一个专注于约鲁巴语(一种非洲声调语言)的数据集,旨在解决声调歧义问题。它包含约鲁巴语长者的访谈语料,每个条目在发声时都进行了声调标注,确保声调标记作为语义数据的一部分,而不是可选的元数据。该数据集用于支持NLP任务,如文本分类和词元分类,特别关注低资源语言和声调语言的处理,以改善下游任务如命名实体识别、机器翻译和语言建模的性能。
The Farahàn Yoruba Elder Corpus is a dataset focused on the Yoruba language, a tonal African language, designed to address tonal disambiguation issues. It contains interview transcripts from Yoruba elders, with each entry annotated for tone at the point of utterance, ensuring tonal marks are integral semantic data rather than optional metadata. This dataset supports NLP tasks such as text classification and token classification, particularly for low-resource and tonal languages, aiming to enhance performance in downstream applications like named entity recognition, machine translation, and language modeling.




