victorachede/tiv-translator-data
收藏资源简介:
English-Tiv Translation Dataset 是一个清理过的英语-Tiv平行语料库,专门用于训练机器翻译模型victorachede/tiv-translator。该数据集包含英语和Tiv(一种尼日利亚语言)的句子对,经过Unicode规范化、轻度质量过滤、基于规范化句子对的去重处理,并通过哈希函数确定性地分割为训练集、验证集和测试集。数据来源包括现有已整理的数据行和从OPUS MT560衍生的数据(遵循CC-BY-4.0许可证)。总共有185,815个被接受的句子对,其中训练集182,128对、验证集1,832对、测试集1,855对,每个句子对包含english、tiv和source字段。需要注意的是,OPUS MT560源数据主要包含宗教/圣经语域的文本,因此为了提升现实世界翻译质量,建议补充由母语者验证的对话、教育、健康、农业、公民和日常领域的数据对。
English-Tiv Translation Dataset is a cleaned English-Tiv parallel corpus specifically designed for training the machine translation model victorachede/tiv-translator. The dataset consists of English and Tiv (a Nigerian language) sentence pairs, which have undergone Unicode normalization, lightweight quality filtering, deduplication based on normalized sentence pairs, and deterministic splitting into training, validation, and test sets via hash functions. Its data sources include existing curated data entries and data derived from OPUS MT560, which is licensed under CC-BY-4.0. In total, there are 185,815 accepted sentence pairs, with 182,128 pairs in the training set, 1,832 pairs in the validation set, and 1,855 pairs in the test set. Each sentence pair contains three fields: english, tiv, and source. It should be noted that the OPUS MT560 source data primarily consists of religious/biblical domain texts. Therefore, to improve real-world translation quality, it is recommended to supplement the corpus with sentence pairs from dialogue, education, healthcare, agriculture, civic, and daily life domains that have been verified by native speakers.




