遇见数据集

Upobhasha: A Multiregional Bengali Dialect Parallel Corpus

收藏
Zenodo2026-09-27 更新2026-10-01 收录
官方服务:

资源简介:

Upobhasha: A Multiregional Bengali Dialect Parallel Corpus is a multilingual and multidialectal Bengali language dataset containing 50,513 parallel records from 10 regional varieties of Bangladesh: Barisal, Chittagong, Kishoreganj, Mymensingh, Narail, Narsingdi, Noakhali, Rangpur, Sylhet, and Tangail. Each record contains a dialect sentence, its Banglish transliteration, a Standard Bangla version, and an English translation, together with regional information and a unique record identifier. The dataset was constructed by integrating and processing material from existing Bengali language resources, including BanglaDial and Vashantor, together with author-curated material. The dataset includes provenance metadata identifying the source and processing history of records. Standard Bangla was prepared by the authors and reviewed by regional validators. Banglish transliterations and English translations were generated using Google Gemini 2.5 Flash and were not subjected to complete record-level human validation. Validation activities focused primarily on the Standard Bangla material and selected regional samples; therefore, the dataset should not be interpreted as having undergone complete manual validation of every field in every record. The dataset is intended for research and development in Bengali dialect processing, multilingual and low-resource natural language processing, dialect identification, machine translation, transliteration, linguistic analysis, and related computational language research. The repository documentation provides the dataset schema, provenance information, validation protocol, validation summary, licensing information, and citation metadata.

提供机构:
Zenodo
创建时间:
2026-09-27
二维码
社区交流群
二维码
科研交流群
商业服务