遇见数据集

LGBTQIAphobia dataset (augmented and balanced)

收藏
Zenodo2025-05-23 更新2026-05-26 收录
官方服务:

资源简介:

Name: LGBTQIAphobia_dataset_augmented_balancedDescription: Labeled dataset with phrases retrieved from different digital sources (X/twitter, Instagram, TikTok) containing diverse messages directed towards the LGBTQIA+ community. It has 1000 phrases classified as {Non-LGBTQIAphobic (0), LGBTQIAphobic (1)} . It is the balanced version of LGBTQIAphobia_dataset_augmented.Language: Spanish Format: CSV (UTF-8)Structure: id; phrase; class {0,1}Purpose: Be used for fine-tuned models that detect language offensive to Spanish or Latin LGBT communities in digital environments.Sources: X/Twitter, Instagram, TikTok, Youtube commentsSize: 20Kb Ethical considerations: This dataset was created strictly for academic and research purposes. We oppose any type of digital violence, in this case, against the LGBTQIA+ community. The person who was the target of the hate speech has been anonymised, and there is no intention to harm them in any way, either them or the person who delivered the speech. We prioritise the protection of the privacy and confidentiality of vulnerable individuals. To safeguard privacy, we carefully remove any identifying details, such as user IDs, phone numbers, and addresses, before sharing the data with our annotators. All the data we collect is from publicly available sources and does not contain any personal or sensitive information that may jeopardise anyone’s privacy. I request researchers to commit to abiding by ethical guidelines so as not to unnecessarily harm individuals.¿How was it created?- Starting recovery of discriminatory phrases for the LGBTQIA+ community from X/Twitter, Instagram, and Tiktok (197 phrases).- Labelling by 3 raters as non-LGBTphobic (0) and LGBTphobic (1).- Text augmentation was applied through backtranslation and random synonym replacement.- Translating to Spanish part of McGiff, J., & Nikolov, N. S. (2024) dataset and was added under licence CC-BY-4.0- To balance the majority class, we applied the undersampling technique.- Finally, we obtained 1000 tagged phrases for version 1.0.2 of LGBTQIAphobia_augmented_balanced Class distribution class instances 0 513 1 487 where class is0: non-lgbtphobic1: lgbtphobic

提供机构:
Zenodo
创建时间:
2025-05-12
二维码
社区交流群
二维码
科研交流群
商业服务