Umsuka English - isiZulu Parallel Corpus
收藏资源简介:
We’ve developed an open-source, high quality isiZulu parallel corpus that comes from a<br> mixture of domains, taking into account both Southern African context and international<br> English context, by using professional translators. We sourced 5000 English sentences,<br> sampled from News Crawl datasets that were translated into isiZulu. Additionally, we<br> translated 5000 isiZulu sentences, sampled from both the NCHLT monolingual corpus and<br> the open-source documents of the UKZN isiZulu National monolingual corpus, into English.<br> From each set, we separated out 1000 patterns as the evaluation dataset. Since isiZulu is<br> highly morphologically complex, we believe that the English-to-isiZulu evaluation set should<br> be translated at least twice, by different translators which will allow us to calculate<br> human-level BLEU score for the dataset. More details in the provided Data Statement



