DigniFy Multimodal Multilingual Hate Speech Datasets
收藏资源简介:
DigniFy Multimodal Multilingual Hate Speech Datasets Data Contents and Provenance Text english_preprocessed_data.csv: Filtered subset of the “Measuring Hate Speech” dataset from Hugging Face (ucberkeley-dlab/measuring-hate-speech). test_hinglish.parquet and train_hinglish.parquet: Train/validation/test splits of the IPD-Text-Hinglish dataset from Hugging Face (dj-dawgs-ipd/IPD-Text-Hinglish). Images We curated and scraped subsets of images depicting various hateful or offensive gestures and symbols. image_symbols.zip: - swastika/ - finger_gun_to_the_head/ - middle_finger/ - cut_throat_gesture/ - slanted_eyes/ Third‑party datasets linked but not hosted: - Roboflow “Nazi Swastika vs. Hindu Swastika” (rm-yoq1a/swast) - Roboflow “Chinese Eyes” Object Detection Dataset by DaeHo (daeho/5-szwgp) - Roboflow "Middle Finger" Dataset by El Hareketi Deneme: (deneme-yz/el-hareketi-deneme) Audio All raw video files were obtained from the HateMM repository. We provide here only the preprocessing files for the spectrograms and features used in the audio model. License This collection is licensed under CC-BY 4.0 for all derivatives. Third‑party assets remain under their original licenses (RF/RM for Getty, CC-BY for Hugging Face, etc.). Citation If you use this dataset, please cite: bibtex@dataset{bhathawala2025dignify, author = {Tirath Bhathawala}, title = {DigniFy Multimodal Multilingual Hate Speech Datasets}, publisher = {Zenodo}, year = {2025}, doi = {10.5281/zenodo.15637274}} For questions or access to raw data not included here, contact the corresponding author.



