SinclairSchneider/tweets_about_german_politicians_jan_feb_2025_reddit_telegram_classified_embed_reduced_clustered
收藏资源简介:
该数据集代表了一个计算框架处理社交媒体政治内容的高级分析阶段。它包含114,876个社交媒体帖子的过滤子集,这些帖子先前被一个提示驱动的推理模型分类为含有操纵性战略叙事(FIMI)。为了在不依赖预定义类别的情况下促进叙事发现,该数据集通过密集向量嵌入、降维投影(UMAP)和基于密度的聚类分配(HDBSCAN)丰富了原始帖子和元数据。这使得研究人员能够探索个别欺骗性帖子如何聚集成更广泛、协调的宣传活动故事情节(例如,“大替换”、“代理战争”或“系统性背叛”)。数据集结构包括新增的列:embeddings(4096维L2归一化向量,使用Qwen3-Embedding-8B模型生成,基于提示映射到修辞操纵空间)、umap_5d(5维UMAP降维向量,优化密度聚类)、umap_2d(2维UMAP向量用于可视化)以及多个min_cluster_size_*列(不同超参数的HDBSCAN聚类ID,其中min_cluster_size_400为最优设置,产生41个不同聚类)。
This dataset represents the advanced analytical stage of a computational framework designed to process political content in social media. It contains a filtered subset of 114,876 social media posts that were previously classified as containing manipulative strategic narratives (FIMI) by a prompt-driven reasoning model. To facilitate narrative discovery without relying on predefined categories, this dataset enriches the raw posts and metadata with dense vector embeddings, dimensionality reduction projections (UMAP), and multiple density-based clustering assignments (HDBSCAN). This allows researchers to explore how individual deceptive posts aggregate into broader, coordinated campaign storylines (e.g., the Great Replacement, the Proxy War, or Systemic Betrayal). The dataset structure includes new columns: embeddings (4096-dimensional L2-normalized vectors generated using the Qwen3-Embedding-8B model, mapped based on rhetorical manipulation via a prompt), umap_5d (5-dimensional UMAP reduced vectors optimized for density-based clustering), umap_2d (2-dimensional UMAP vectors for visualization), and multiple min_cluster_size_* columns (HDBSCAN cluster IDs with various hyperparameters, where min_cluster_size_400 is optimal, yielding 41 distinct clusters).



