遇见数据集

BhasaBodh: Bridging Bangla Dialects and Romanized Forms through Machine Translation

收藏
NIAID Data Ecosystem2026-05-10 收录
官方服务:

资源简介:

While machine translation has made significant strides for high-resource languages, many regional languages and their dialects, such as the Bangla variants Chittagong and Sylhet, remain underserved. Existing resources are often insufficient for robust sentence-level evaluation and overlook the widespread real-world practice of romanization, the common practice of typing native languages using the Latin script in digital communication. To address these gaps, we introduce BhasaBodh, a comprehensive benchmark for Bangla dialectal machine translation. We construct and release a sentence-level parallel dataset for Chittagong and Sylhet dialects aligned with Standard Bangla and English, create a novel romanized version of the dialectal data to facilitate evaluation in realistic multi-script scenarios, and provide the first comprehensive performance baselines by fine-tuning two powerful multilingual models, NLLB-200 and mBART-50, on seven distinct translation tasks. However, complex cross-lingual and cross-script translation remains a significant challenge. BhasaBodh lays the groundwork for future research in low-resource dialectal NLP, offering a valuable resource for developing more inclusive and practical translation systems. Citation: Bhuiyan, Md Tofael Ahmed, Md Abdur Rahman, and Abdul Kadar Muhammad Masum. "BhasaBodh: Bridging Bangla Dialects and Romanized Forms through Machine Translation." In Proceedings of the Second Workshop on Bangla Language Processing (BLP-2025), pp. 113-118. 2025.

创建时间:
2026-03-18
二维码
社区交流群
二维码
科研交流群
商业服务