CyberbullyX-63K: A Large-Scale Twitter/X Multilingual Hinglish Tweets Dataset for Cyberbullying Detection
收藏资源简介:
CyberBullyX-63K is a large-scale Twitter/X dataset developed for cyberbullying detection and harmful content classification research. The dataset contains 63,145 publicly available tweets collected through the official Twitter/X Developer API between January 2020 and May 2026. Data collection was conducted in multiple phases across different time periods to ensure temporal diversity and broad representation of online interactions, abusive language, toxic communication, harassment, and non-cyberbullying content. The dataset was created to support research in cyberbullying detection using machine learning, deep learning, and large language model (LLM)-based approaches. In the final_label column, 0 represents not cyberbullying whereas 1 represents the text is classified as cyberbullying. CyberBullyX-63K can be used for cyberbullying detection, abusive language identification, online harassment analysis, automated moderation systems, comparative evaluation of LLMs, and NLP classification research. All tweets were collected in compliance with Twitter/X developer policies and include only publicly accessible content intended for academic and non-commercial research purposes.



