遇见数据集

Hate_annotated_data_Bhojpuri_hindi_english

收藏
Zenodo2026-03-09 更新2026-05-26 收录
官方服务:

资源简介:

This is a multilingual dataset that contains about 7000 comments extracted from Social media platforms . The comments are recorded in three languages: Hindi, English and Bhojpuri(a dialect of Hindi). The samples are in both Roman as well as Devanagari scripts. There are two classes to identify the samples, namely O(offensive) and NO(not offensive). Any comment that incites hate, uses objectionable words or is targeted against any community is considered as O(objectionable), the others are NO(non objectionable). Annotators were hired to annotate the data to the best case.

本数据集为多语言数据集,收录了从社交媒体平台提取的约7000条评论。该评论涵盖三种语言:印地语(Hindi)、英语以及博杰普尔语(Bhojpuri,印地语方言之一)。所有样本同时采用罗马字母(Roman)与天城文(Devanagari)两种书写体系。本数据集设有两类分类标签:O(攻击性,offensive)与NO(非攻击性,not offensive)。其中,任何煽动仇恨、使用冒犯性词汇或针对任何社群的评论均被归类为O类(冒犯性),其余评论则归为NO类(非冒犯性)。本数据集由雇佣的标注人员以最优标注标准完成全部标注工作。

提供机构:
Zenodo
创建时间:
2025-11-18
二维码
社区交流群
二维码
科研交流群
商业服务