遇见数据集

BenCyber: A Multi-Aspect Bengali Cyberbullying Corpus with Demographics and Social Engagement Metrics

收藏
Mendeley Data2026-08-04 收录
官方服务:

资源简介:

This dataset features 44,001 meticulously annotated Bangla social media comment records designed for granular cyberbullying, hate speech, and online harassment detection. The corpus captures contextual linguistic features by tracking five distinct target profiles (Actors, Social Figures, Singers, Politicians, and Sports Personalities) along with their gender identification. Each comment record includes metadata on social engagement through public reaction numbers and is manually categorized into one of five operational classes: Neutral, Harassment, Sexual Aggression, Hate Speech, and Violent Extremism. This structured contextual layout is highly optimized for developing explainable machine learning architectures, deep learning classifiers (such as Banglish/Bangla BERT, RoBERTa, and BiLSTM), and multi-task learning frameworks aimed at protecting vulnerable digital demographics in low-resource language environments.

本数据集包含44001条经严谨标注的孟加拉语(Bangla)社交媒体评论记录,专为细粒度网络欺凌、仇恨言论与在线骚扰检测任务打造。该语料库通过追踪五类明确的目标对象画像——演员、社会名流、歌手、政客与体育名人——及其性别身份,捕捉语境化语言特征。每条评论记录均附带包含公开互动数值的社交元数据,并被人工划分为以下五类任务类别之一:中性、骚扰、性侵犯、仇恨言论与暴力极端主义。这种结构化语境化设计高度适配可解释机器学习架构、深度学习分类器(如孟加拉拉丁化语(Banglish)/孟加拉语BERT、RoBERTa及BiLSTM)以及多任务学习框架的开发需求,旨在保护低资源语言环境中的弱势数字群体。

创建时间:
2026-07-08
二维码
社区交流群
二维码
科研交流群
商业服务