遇见数据集

YouTube-BN

收藏
Zenodo2026-08-08 更新2026-08-20 收录
官方服务:

资源简介:

youtube_bn_topic is a Bengali-language comment dataset developed for six-way topic classification and analysis of user-generated social media content. The dataset primarily consists of Bengali comments collected from YouTube, with additional Bengali comments manually collected from other social media platforms, including Facebook, Twitter/X, and Instagram, where relevant Bengali bullying and abusive content was identified. The majority of the dataset is sourced from YouTube. Approximately 140,000 comments were initially collected from around 410 YouTube video IDs using the YouTube Data API v3. The collected comments subsequently underwent extensive cleaning and preprocessing to remove noisy, duplicate, irrelevant, and unsuitable instances. After the preprocessing and quality-control process, the final dataset contains 67,982 Bengali comments. All candidate comments were annotated using a six-way labeling scheme: Geopolitical, Personal, Political, Religious, Gender Abusive, and Non-abusive. Comments were first labeled by one author and subsequently reviewed and refined by a second annotator. Cases involving disagreements were resolved through discussion and consensus to determine the final label. Since the second annotator reviewed and refined the initial annotations rather than independently re-labeling the samples, Cohen's κ was not applicable under this annotation protocol. Independent double-coding on a held-out subset is identified as a potential direction for a future dataset release. The final dataset contains 32,938 Non-abusive, 9,155 Religious, 9,057 Personal, 7,006 Political, 5,746 Geopolitical, and 4,080 Gender Abusive comments. The dataset is intended to support research in Bengali natural language processing (NLP), topic classification, abusive-language analysis, social media text mining, and computational analysis of Bengali user-generated content.

提供机构:
Zenodo
创建时间:
2026-08-08
二维码
社区交流群
二维码
科研交流群
商业服务