遇见数据集

SinclairSchneider/tweets_about_german_politicians_jan_feb_2025_reddit_and_telegram_classified

收藏
Hugging Face2026-05-16 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是用于检测操纵性政治叙事的文本分类数据集,属于政治科学、社交媒体和情感分析领域。数据集包含1,255,895条未经过滤的社交媒体帖子,收集自X(前Twitter)、Reddit和Telegram平台,语言分布约为80%德语和20%英语。数据收集于2025年1月至2月,内容涉及德国政治家(如Alice Weidel、Karl Lauterbach和现任德国总理Friedrich Merz)的讨论。与原始数据不同,该数据集通过基于提示的分类处理,使用Qwen3.5-122B-A10B-FP8模型检测外国信息操纵和干扰(FIMI),以分离协调操纵性叙事与合法政治批评。数据集新增了分类列(classified、contains_narrative和reasoning),支持文本分类和情感分析任务。

This dataset represents a critical stage in a broader framework for processing political content in social media ecosystems. It comprises an unfiltered collection of 1,255,895 short social media posts collected from X (formerly Twitter), Reddit, and Telegram. The language distribution is approximately 80% German and 20% English. The data was collected between January and February 2025, capturing discourse surrounding German politicians, including prominent figures like Alice Weidel, Karl Lauterbach, and the current German Chancellor, Friedrich Merz. Unlike raw scrapes, this dataset includes a specialized prompt-based filtering classification to detect Foreign Information Manipulation and Interference (FIMI), separating coordinated manipulative storylines from legitimate political critique. It supports text-classification and sentiment-analysis tasks, with added classification columns (classified, contains_narrative, reasoning) processed using the Qwen3.5-122B-A10B-FP8 model.

提供机构:
SinclairSchneider
二维码
社区交流群
二维码
科研交流群
商业服务