Hybrid Hinglish Code-Mixed Dataset for Language Identification and Text Normalization
收藏资源简介:
This dataset is a hybrid code-mixed Hinglish corpus developed for research in Language Identification (LI) and Text Normalization. It consists of 968,231 annotated tokens collected and curated from publicly available code-mixed resources. Each record contains the original token, its initial language label, the normalized form (where applicable), and the final refined language label. The dataset includes English words, Romanized Hindi words, slang expressions, code-mixed tokens, acronyms, and undefined or out-of-vocabulary (OOV) words. It is intended for developing and evaluating machine learning and deep learning models for low-resource multilingual NLP tasks such as language identification, lexical normalization, code-mixed text processing, and machine translation.



