遇见数据集

EN-UR Parallel Dataset

收藏
Mendeley Data2024-01-31 更新2024-06-26 收录
官方服务:

资源简介:

The parallel corpus is gathered from various accessible sources, encompassing distinct domains such as Journalism, the Quran, the Bible, News, Subtitles, Movies, COVID-19, and Human Rights. The gathered data went through a thorough preprocessing pipeline designed to guarantee the utmost quality for training translation models.

本平行语料库(parallel corpus)采集自各类公开可获取的来源,涵盖新闻业、古兰经、圣经、新闻、字幕、电影、新冠疫情(COVID-19)与人权等多个不同领域。所采集的数据经过一套全面严谨的预处理流水线,以确保用于翻译模型训练的语料达到最高质量水准。

创建时间:
2024-01-31
二维码
社区交流群
二维码
科研交流群
商业服务