BSC-LT/BSC_ParaMT_8
收藏资源简介:
BSC_ParaMT_8是一个大规模多语言平行语料库,主要覆盖加泰罗尼亚语、西班牙语和英语与阿拉伯语、印地语、中文、日语、韩语之间的配对。该数据集通过聚合和精心过滤多个公共来源(如Tatoeba、UNPC、NLLB、WikiMatrix等)构建,提供句子级别的对齐数据,用于训练机器翻译系统。其中,西班牙语部分包含通过将原始英语句子翻译成西班牙语生成的合成数据;加泰罗尼亚语部分包含通过将原始英语和西班牙语句子翻译成加泰罗尼亚语生成的合成数据,这些合成文本使用了SalamandraTA 7B Instruct模型生成。数据集旨在促进机器翻译研究,支持多语言NLP应用,并改善对代表性不足语言对的翻译可访问性。
BSC_ParaMT_8 is a large-scale multilingual parallel corpus primarily covering parallel pairs between Catalan, Spanish, English and Arabic, Hindi, Chinese, Japanese, Korean. Constructed by aggregating and meticulously filtering multiple public resources such as Tatoeba, UNPC, NLLB, WikiMatrix, etc., it provides sentence-level aligned data for training machine translation systems. Specifically, the Spanish subset includes synthetic data generated by translating original English sentences into Spanish; the Catalan subset contains synthetic data produced by translating original English and Spanish sentences into Catalan, where the synthetic texts were generated using the SalamandraTA 7B Instruct model. This dataset aims to advance machine translation research, support multilingual natural language processing (NLP) applications, and improve translation accessibility for underrepresented language pairs.




