遇见数据集

Hate_annotated_data_Bhojpuri_hindi_english

收藏
Zenodo2026-05-19 更新2026-05-26 收录
官方服务:

资源简介:

This is a multilingual dataset that contains about 7000 comments extracted from Social media platforms . The comments are recorded in three languages: Hindi, English and Bhojpuri(a dialect of Hindi). The samples are in both Roman as well as Devanagari scripts. There are two classes to identify the samples, namely O(offensive) and NO(not offensive). Any comment that incites hate, uses objectionable words or is targeted against any community is considered as O(objectionable), the others are NO(non objectionable). Annotators were hired to annotate the data to the best case.

提供机构:
Zenodo
创建时间:
2025-11-07
二维码
社区交流群
二维码
科研交流群
商业服务