Parallel Corpus of Translationese
收藏资源简介:
本数据集名为‘Parallel Corpus of Translationese’,由海法大学计算机科学系和萨尔兰大学计算语言学系共同创建。数据集包含英法和英德双语平行语料,涵盖议会进程、文学作品、TED演讲转录及政治评论等多种文本类型,总计约347,000条。数据集经过严格预处理,确保每条数据的一对一句子对齐,适用于翻译学研究,特别是翻译方向识别。该数据集旨在解决翻译文本的自动识别问题,支持机器翻译和人类翻译研究,对翻译学领域具有重要意义。
This dataset, named *Parallel Corpus of Translationese*, was jointly developed by the Department of Computer Science at the University of Haifa and the Department of Computational Linguistics at Saarland University. It comprises bilingual parallel corpora for English-French and English-German language pairs, covering a wide range of text types including parliamentary proceedings, literary works, TED talk transcripts, and political commentaries, with a total of approximately 347,000 sentence pairs. The dataset has undergone rigorous preprocessing to guarantee one-to-one sentence alignment for every entry, rendering it suitable for translation studies, especially research on translation direction identification. This dataset is designed to tackle the problem of automatic identification of translated texts, supporting research on both machine translation and human translation, and carries substantial significance for the field of translation studies.




