遇见数据集

English-Spanish website parallel corpus (Processed)

收藏
data.europa2024-06-27 收录
官方服务:

资源简介:

This is a parallel corpus of bilingual texts crawled from multilingual websites, which contains 909 TUs. Manual validation has been performed on a sample of the data. This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) actions SMART 2014/1074 and SMART 2015/1091. For further information on the project: http://lr-coordination.eu.

本数据集为从多语言网站爬取获得的双语平行语料库,共计包含909个翻译单元(Translation Unit,TU)。已对该数据集的部分样本完成人工校验工作。 本数据集依托欧洲语言资源协调(European Language Resource Coordination, ELRC)框架下的「欧洲连接设施-自动化翻译(CEF.AT)」项目行动SMART 2014/1074与SMART 2015/1091开发研制。如需了解该项目的更多相关信息,请访问:http://lr-coordination.eu。

二维码
社区交流群
二维码
科研交流群
商业服务