DWADASH
收藏资源简介:
DWADASH是目前规模最大的孟加拉语多方言平行语料库,由锡尔赫特工程学院构建,旨在为多方言神经机器翻译提供统一基准。该语料库包含14,562行对齐数据,覆盖12种方言区域,共51,541个非空平行句对,通过整合Ancholik-NER、Anubhuti等七项现有数据集并额外添加2,510条经母语专家验证的平行句对而成。构建过程包括数据集成、去重、Unicode标准化及分词对齐等步骤,最终按80/10/10划分为训练、开发与测试集。该数据集主要应用于多方向方言翻译任务(方言→标准语、标准语→方言、方言间直接翻译),旨在克服传统中继翻译的误差累积问题,为方言使用者提供包容性的数字语言服务。
DWADASH is the largest Bengali multi-dialect parallel corpus to date, constructed by Sylhet Engineering College, with the objective of providing a unified benchmark for multi-dialect neural machine translation. This corpus consists of 14,562 aligned data entries, spanning 12 dialect regions, and contains a total of 51,541 non-empty parallel sentence pairs. It is compiled by integrating seven existing datasets such as Ancholik-NER and Anubhuti, along with an additional 2,510 parallel sentence pairs validated by native language experts. The construction process comprises data integration, deduplication, Unicode standardization, word segmentation and alignment, and finally splits the corpus into training, development and test sets with a ratio of 80:10:10. This dataset is primarily applied to multi-directional dialect translation tasks (dialect → Standard Bengali, Standard Bengali → dialect, and direct cross-dialect translation), aiming to resolve the error accumulation issue in traditional relay translation and provide inclusive digital language services for dialect speakers.

- 1Unified Multi-Dialectal Neural Machine Translation for Bangla Using the Dwadash Benchmark Corpus锡尔赫特工程学院 · 2026年



