遇见数据集

Mute Cods: a Multilingual Telegram Conspiracy DataSet

收藏
Zenodo2026-03-03 更新2026-05-26 收录
官方服务:

资源简介:

This is the dataset presented in the paper 'Mute Cods: A multilingual Telegram Dataset with Benchmark Models for Conspiracy Theory Detection', accepted at LREC conference 2026, by Laken, Marino, Piot, Bassi, Fomsgaard, Maggini, Vieira, García and Tonelli. The dataset consists of 5750 messages across English, Dutch, Italian, Spanish and Portuguese from 87 channels documented as disseminating conspiracist and extremist content. Domain experts annotated messages for conspiracist tone, population replacement conspiracy theories, vaccine conspiracies, and hate speech. This dataset is pseudonomyzed. We release both the raw annotations and the aggregated labels with train/test split used to train the models as reported in the paper. As this is social media user data, we request users not to share the dataset with third parties; rather, send them the link to our repository, and we will grant them access as well. If you use this dataset, please cite our paper: Katarina Laken, Erik Bran Marino, Paloma Piot, Davide Bassi, Søren Fomsgaard, Michele Maggini, Renata Vieira, Marcos García, and Sara Tonelli. 2026. Mute Cods: A Multilingual Telegram Dataset with Benchmark Models for Conspiracy Theory Detection. In Proceedings of the Fifteenth International Conference on Language Resources and Evaluation (LREC 2026). Palma de Mallorca, Spain.

提供机构:
Zenodo
创建时间:
2026-03-03
二维码
社区交流群
二维码
科研交流群
商业服务