遇见数据集

Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset

收藏
Zenodo2026-08-12 更新2026-08-13 收录
官方服务:

资源简介:

This dataset addresses an empirically confirmed gap in the code-mixed sentiment analysis literature: a systematic literature review of 48 retrieved records (39 core studies) found zero existing Gujarati-English code-mixed sentiment datasets, despite Gujarati having over 55 million native speakers. This dataset provides the first Gujarati-English resource of this kind, alongside Hindi-English, Gujarati-Hindi, and trilingual code-mixed comments collected from the same corpus. Key statistics:- 21,729 deduplicated, language-identified, code-mixed YouTube comments (Gujarati-English, Hindi-English, Gujarati-Hindi, or trilingual)- 600 comments individually sentiment-labeled via LLM-assisted annotation (gold subset)- 21,129 comments sentiment-labeled via a semi-supervised classical ML model trained on the gold subset (silver subset)- Full Code-Mixing Index (CMI) score computed for every comment- Every comment traceable to its source video and collection channel

提供机构:
Zenodo
创建时间:
2026-08-12
二维码
社区交流群
二维码
科研交流群
商业服务