遇见数据集

Spanish Gender-Related Cyberbullying Dataset

收藏
Zenodo2026-07-17 更新2026-08-01 收录
官方服务:

资源简介:

The Spanish Gender-Related Cyberbullying Dataset contains 2,548 original Spanish-language tweets collected from public Twitter posts between 23 February and 3 April 2023 using Twint. Five queries were used to retrieve content related to public debates on sexual-consent legislation, public figures linked to equality policies, derogatory labels targeting feminist positions, anti-feminist expressions, and references to feminist groups or movements. Only original tweets were retained. Duplicate tweets, non-Spanish content, messages without textual content, and tweets unrelated to gender-related cyberbullying were excluded. Relevance was assessed through manual review, and all remaining valid tweets were included in the corpus. The corpus was annotated by three annotators using three levels of offensiveness. Disagreements were resolved through majority voting. The final distribution comprises 1,176 non-offensive tweets, 769 moderately offensive tweets, and 603 offensive tweets. For binary classification, the moderately offensive and offensive categories can be merged into a single offensive class. The dataset supports research on gender-related cyberbullying detection, offensive-language classification, linguistic metadata, and Natural Language Processing for Spanish social-media text. It accompanies the paper “Semantic and Linguistic Metadata Fusion for Gender-Related Cyberbullying Detection”.

提供机构:
Zenodo
创建时间:
2026-07-17
二维码
社区交流群
二维码
科研交流群
商业服务