Mute Cods: a Multilingual Telegram Conspiracy DataSet
收藏资源简介:
This is the dataset presented in the paper 'Mute Cods: A multilingual Telegram Dataset with Benchmark Models for Conspiracy Theory Detection', accepted at LREC conference 2026, by Laken, Marino, Piot, Bassi, Fomsgaard, Maggini, Vieira, García and Tonelli. The dataset consists of 5750 messages across English, Dutch, Italian, Spanish and Portuguese from 87 channels documented as disseminating conspiracist and extremist content. Domain experts annotated messages for conspiracist tone, population replacement conspiracy theories, vaccine conspiracies, and hate speech. This dataset is pseudonomyzed. We release both the raw annotations and the aggregated labels with train/test split used to train the models as reported in the paper. As this is social media user data, we request users not to share the dataset with third parties; rather, send them the link to our repository, and we will grant them access as well. If you use this dataset, please cite our paper: Katarina Laken, Erik Bran Marino, Paloma Piot, Davide Bassi, Søren Fomsgaard, Michele Maggini, Renata Vieira, Marcos García, and Sara Tonelli. 2026. Mute Cods: A Multilingual Telegram Dataset with Benchmark Models for Conspiracy Theory Detection. In Proceedings of the Fifteenth International Conference on Language Resources and Evaluation (LREC 2026). Palma de Mallorca, Spain.



