English-Tamil-Parallel-Corpus
收藏资源简介:
由Moratuwa大学的National Languages Processing Center准备的英泰平行语料库。数据已经过清洗和校准。数据来源于公开的政府资源,如年度报告、采购报告、通函和网站。每个word/pdf文件被转换为文本文件,并使用自定义工具修复了unicode错误。泰米尔语和英语文件经过手动句子对齐,所有拼写和语法错误都经过手动修正。
The English-Tamil parallel corpus prepared by the National Languages Processing Center at the University of Moratuwa. The data has been cleaned and calibrated. The data originates from publicly available government resources, such as annual reports, procurement reports, circulars, and websites. Each word/pdf file was converted into a text file, and Unicode errors were rectified using custom tools. The Tamil and English files were manually sentence-aligned, and all spelling and grammatical errors were manually corrected.
数据集概述
数据集名称
English-Tamil parallel Corpus
数据集准备机构
National Languages Processing Center, University of Moratuwa
数据集内容
- En-Ta Glossary Line Count: 22477
- En-Ta Corpus Line Count: 8950
数据来源
数据提取自公开的政府资源,包括年度报告、采购报告、通知和网站。
数据处理
- 每个word/pdf文件转换为文本文件。
- 使用定制工具修复unicode错误。
- 手动进行泰米尔语和英语文件的句子对齐。
- 手动修正所有拼写和语法错误。
引用信息
若使用此数据集,请引用以下出版物: Fernando, A., Ranathunga, S., & Dias, G. (2020). Data Augmentation and Terminology Integration for Domain-Specific Sinhala-English-Tamil Statistical Machine Translation. arXiv preprint arXiv:2011.02821.




