tg-moises: A Multilingual Telegram Corpus Annotated for Malicious-Content Taxonomy (MOISES Project)
收藏资源简介:
This dataset contains 6,805,925 messages collected from 490 public Telegram channels associated with conspiracy theories, anti-vaccine movements, and far-right/nationalist narratives across multiple languages (including English, Spanish, Turkish, Lithuanian, French and German). It supports the MOISES line of research on profiling malicious actors and disinformation dynamics in online social networks. A subset of 20,941 messages (466 channels) includes human annotations using a five-dimension taxonomy: Role, Tactic, Feature, Target, and Vulnerability. Based on the related preprint, taxonomy construction combined subject-matter expert workshops and literature review, and was applied in Telegram case-study analyses. From the related preprint: research activities were conducted in 2023; annotation was performed with multi-annotator procedures; and the annotation infrastructure was configured with privacy-preserving choices for handling public channel data. Funding: MARTINI project (Malicious Actors Profiling and Detection in Online Social Networks through Artificial Intelligence), CHIST-ERA call, grant CHIST-ERA-21-OSNEM-004.



