遇见数据集

kahakashanashraf/BAAD-Dataset: BAAD-Dataset

收藏
Zenodo2026-03-08 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains 46,128 labeled Bengali comments for the task of arrogance detection. While existing datasets focus heavily on hate speech or cyberbullying, this dataset addresses the subtle linguistic nuances of "arrogance", characterized by overbearing pride, lack of empathy, and social superiority, which is often expressed without overt toxicity. Dataset Structure: The dataset is provided in a single .csv file with the following columns:comment: The raw Bengali text.source: The origin of the comment (online or AI).weak_label: Initial label assigned by heuristic functions.snorkel_label: Refined label produced by the Snorkel framework.final_label: The target label for classification.1: Arrogant0:Non-arrogant**further a automaited English translated dataset is attached as test_translated_data.csv

提供机构:
Zenodo
创建时间:
2026-03-08
二维码
社区交流群
二维码
科研交流群
商业服务