遇见数据集

Hybrid Hinglish Code-Mixed Dataset for Language Identification and Text Normalization

收藏
Zenodo2026-07-15 更新2026-08-02 收录
官方服务:

资源简介:

This dataset is a hybrid code-mixed Hinglish corpus developed for research in Language Identification (LI) and Text Normalization. It consists of 968,231 annotated tokens collected and curated from publicly available code-mixed resources. Each record contains the original token, its initial language label, the normalized form (where applicable), and the final refined language label. The dataset includes English words, Romanized Hindi words, slang expressions, code-mixed tokens, acronyms, and undefined or out-of-vocabulary (OOV) words. It is intended for developing and evaluating machine learning and deep learning models for low-resource multilingual NLP tasks such as language identification, lexical normalization, code-mixed text processing, and machine translation.

提供机构:
Zenodo
创建时间:
2026-07-15
二维码
社区交流群
二维码
科研交流群
商业服务