YouTube and TikTok评论数据集
收藏资源简介:
该数据集由国际巴尔干大学的研究人员构建,包含来自YouTube和TikTok平台的4500条评论,涵盖塞尔维亚、克罗地亚和波斯尼亚三种语言。数据集涉及音乐、政治、体育、模特、影响者、性别主义和一般话题等多个类别,旨在用于评估大型语言模型在低资源环境下的毒性语言检测能力。数据集已进行手动标注,标注过程由两位作者完成,并采用Cohen's Kappa系数进行可靠性评估。
This dataset was constructed by researchers from the International Balkan University, containing 4,500 comments sourced from YouTube and TikTok platforms across three languages: Serbian, Croatian, and Bosnian. It covers multiple categories including music, politics, sports, modeling, influencers, sexism, and general topics, and is designed to evaluate the toxic language detection capabilities of large language models (LLMs) in low-resource environments. The dataset has undergone manual annotation completed by two authors, and its annotation reliability was assessed using the Cohen's Kappa coefficient.

- 1Large Language Models for Toxic Language Detection in Low-Resource Balkan Languages国际巴尔干大学 · 2025年



