EN-UR Parallel Dataset
收藏官方服务:
资源简介:
The parallel corpus is gathered from various accessible sources, encompassing distinct domains such as Journalism, the Quran, the Bible, News, Subtitles, Movies, COVID-19, and Human Rights. The gathered data went through a thorough preprocessing pipeline designed to guarantee the utmost quality for training translation models.
本平行语料库(parallel corpus)采集自各类公开可获取的来源,涵盖新闻业、古兰经、圣经、新闻、字幕、电影、新冠疫情(COVID-19)与人权等多个不同领域。所采集的数据经过一套全面严谨的预处理流水线,以确保用于翻译模型训练的语料达到最高质量水准。
创建时间:
2024-01-31



