Hinglish–English Parallel Corpus for Code-Mixed Machine Translation (111,133 Sentence Pairs)
收藏官方服务:
资源简介:
This dataset contains 111,133 parallel Hinglish–English sentence pairs developed for research in code-mixed natural language processing and machine translation. The source sentences are written in Romanized Hinglish (Hindi-English code-mixed text), while the target sentences are their corresponding English translations. The corpus is suitable for training and evaluating machine translation systems, large language models, text normalization techniques, language identification, transliteration, and other NLP applications involving multilingual and code-mixed text. The dataset can also be used for benchmarking transformer-based translation models such as mT5, mBART, and T5 variants.
提供机构:
Zenodo创建时间:
2026-07-16



