遇见数据集

Replication Data for: Hindi-English code-mixed Twitter dataset

收藏
Harvard Dataverse2024-02-06 更新2026-04-09 收录
官方服务:

资源简介:

This directory contains a large-scale Hindi-English code-mixed corpus collected from Twitter between 2010-2022. We have removed the identifiers for anonymizing the dataset. We have de-anonymized the tweet author ids. Additionally, we have calculated code-mixing index (CMI) and the language of the texts (Hindi, English or, Hindi-English code-mixed).

创建时间:
2023-01-01
二维码
社区交流群
二维码
科研交流群
商业服务