遇见数据集

A Multilingual Telegram Corpus Annotated for Malicious-Content Taxonomy (MOISES Project)

收藏
Zenodo2026-07-13 更新2026-08-02 收录
官方服务:

资源简介:

This dataset contains 6,805,925 messages collected from 490 public Telegram channels associated with conspiracy theories, anti-vaccine movements, and far-right/nationalist narratives across multiple languages (including English, Spanish, Turkish, Lithuanian, French and German). It supports the MOISES line of research on profiling malicious actors and disinformation dynamics in online social networks. A subset of 20,941 messages (466 channels) includes human annotations using a five-dimension taxonomy: Role, Tactic, Feature, Target, and Vulnerability. Based on the related preprint, taxonomy construction combined subject-matter expert workshops and literature review, and was applied in Telegram case-study analyses. From the related preprint: research activities were conducted in 2023; annotation was performed with multi-annotator procedures; and the annotation infrastructure was configured with privacy-preserving choices for handling public channel data. Funding: MARTINI project (Malicious Actors Profiling and Detection in Online Social Networks through Artificial Intelligence), CHIST-ERA call, grant PCI2022-134990-2. Files included full_corpus.jsonl.gz: Full corpus export with one JSON object per line; includes message-level fields and the channel source channel label. human_labeled_subset.jsonl.gz: Human-labeled subset where taxonomy_human contains at least one annotation in the taxonomy dimensions. channel_index.csv: Per-channel index with message counts and whether the channel contains human-labeled messages. channel_metadata.json: Channel-level aggregate metadata exported from the _META collection (e.g., language distribution and topic-model outputs when available). README.md: Data dictionary and usage notes describing fields, structure, and export context.

提供机构:
Zenodo
创建时间:
2026-07-13
二维码
社区交流群
二维码
科研交流群
商业服务