KU-Bangla: A Comprehensive Bangla Dataset for Language Modeling and Translation
收藏资源简介:
This dataset focuses on Bangla-to-English translation, compiled from diverse sources including textbooks, dictionaries, websites, and AI-generated content. Consisting of about 6000 entries, the dataset covers mostly Standard Colloquial Bangla (Chalit bhasha), and non-literal idioms or phrases have their keywords’ meanings and tenses defined. The dataset is meant for the translation process of Bangla to English using LLM concepts. It is designed to train and evaluate machine translation models on Bangla-to-English. The dataset consists of four columns. The first column contains the Bangla sentence. The second column contains grammatical labels identifying specific parts of speech (e.g., nouns, pronouns, verbs, and adjectives) present in the Bangla sentence. The third column contains the tense label of the Bangla sentence. The fourth column contains the corresponding English translation of the Bangla sentence.



