遇见数据集

Or4kool/english-south-african-languages

收藏
Hugging Face2026-05-23 更新2026-05-31 收录
官方服务:

资源简介:

这是一个多语言平行语料数据集,专门用于机器翻译任务。数据集包含英语与多种非洲语言之间的双向翻译对,涉及的语言包括南非荷兰语(afr)、北索托语(nbl)、科萨语(xho)、祖鲁语(zul)、南索托语(nso)、茨瓦纳语(tsn)、文达语(ven)和聪加语(tso)。数据集中每个样本都包含源文本、目标文本、源语言、目标语言、语言对和任务类型等特征字段。数据集总大小约为2.77 GB,包含约498万个样本,分为多个语言对分片,每个分片提供双向翻译数据(如en_afr和afr_eng),适用于训练和评估多语言机器翻译模型。

This is a multilingual parallel corpus dataset specifically designed for machine translation tasks. The dataset contains bidirectional translation pairs between English and multiple African languages, including Afrikaans (afr), Northern Sotho (nbl), Xhosa (xho), Zulu (zul), Southern Sotho (nso), Tswana (tsn), Venda (ven), and Tsonga (tso). Each sample in the dataset includes feature fields such as source text, target text, source language, target language, language pair, and task type. The total size of the dataset is approximately 2.77 GB, containing about 4.98 million samples. It is divided into multiple language pair shards, each providing bidirectional translation data (such as en_afr and afr_eng), which is suitable for training and evaluating multilingual machine translation models.

提供机构:
Or4kool
二维码
社区交流群
二维码
科研交流群
商业服务