Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset
收藏资源简介:
This dataset addresses an empirically confirmed gap in the code-mixed sentiment analysis literature: a systematic literature review of 48 retrieved records (39 core studies) found zero existing Gujarati-English code-mixed sentiment datasets, despite Gujarati having over 55 million native speakers. This dataset provides the first Gujarati-English resource of this kind, alongside Hindi-English, Gujarati-Hindi, and trilingual code-mixed comments collected from the same corpus. Key statistics:- 21,729 deduplicated, language-identified, code-mixed YouTube comments (Gujarati-English, Hindi-English, Gujarati-Hindi, or trilingual)- 600 comments individually sentiment-labeled via LLM-assisted annotation (gold subset)- 21,129 comments sentiment-labeled via a semi-supervised classical ML model trained on the gold subset (silver subset)- Full Code-Mixing Index (CMI) score computed for every comment- Every comment traceable to its source video and collection channel



