BPCC Parallel Corpus
收藏资源简介:
BPCC Parallel Corpus是由语言技术研究中心(LTRC)在IIIT-海得拉巴创建的一个大规模平行语料库,包含2.3亿个句子对,涵盖英语和22种印度语言。该数据集通过手动验证的子集BPCC-Human Corpus提供了高质量的数据,适用于训练和评估机器翻译模型。数据集的创建旨在解决印度语言在机器翻译中的独特挑战,如复杂的形态结构和多样化的脚本使用。该数据集的应用领域广泛,包括跨语言通信、教育和医疗等,旨在促进印度多语言生态系统的发展。
The BPCC Parallel Corpus is a large-scale parallel corpus developed by the Language Technology Research Center (LTRC) at IIIT-Hyderabad, containing 230 million sentence pairs covering English and 22 Indian languages. A manually validated subset of this corpus, the BPCC-Human Corpus, provides high-quality data suitable for training and evaluating machine translation models. The development of this corpus aims to address the unique challenges in machine translation for Indian languages, such as complex morphological structures and diverse script systems. This corpus has broad application areas including cross-lingual communication, education, healthcare and more, with the goal of promoting the development of India's multilingual ecosystem.

- 1BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages语言技术研究中心,IIIT-海得拉巴,海得拉巴,特伦甘纳邦,印度 · 2024年



