Users data points.
收藏资源简介:
Social networks are a battlefield for political propaganda. Protected by the anonymity of the internet, political actors use computational propaganda to influence the masses. Their methods include the use of synchronized or individual bots, multiple accounts operated by one social media management tool, or different manipulations of search engines and social network algorithms, all aiming to promote their ideology. While computational propaganda influences modern society, it is hard to measure or detect it. Furthermore, with the recent exponential growth in large language models (L.L.M), and the growing concerns about information overload, which makes the alternative truth spheres more noisy than ever before, the complexity and magnitude of computational propaganda is also expected to increase, making their detection even harder. Propaganda in social networks is disguised as legitimate news sent from authentic users. It smartly blended real users with fake accounts. We seek here to detect efforts to manipulate the spread of information in social networks, by one of the fundamental macro-scale properties of rhetoric—repetitiveness. We use 16 data sets of a total size of 13 GB, 10 related to political topics and 6 related to non-political ones (large-scale disasters), each ranging from tens of thousands to a few million of tweets. We compare them and identify statistical and network properties that distinguish between these two types of information cascades. These features are based on both the repetition distribution of hashtags and the mentions of users, as well as the network structure. Together, they enable us to distinguish (p − value = 0.0001) between the two different classes of information cascades. In addition to constructing a bipartite graph connecting words and tweets to each cascade, we develop a quantitative measure and show how it can be used to distinguish between political and non-political discussions. Our method is indifferent to the cascade’s country of origin, language, or cultural background since it is only based on the statistical properties of repetitiveness and the word appearance in tweets bipartite network structures.
社交网络已然成为政治宣传的角力场。依托互联网的匿名性庇护,政治行为主体借助计算宣传手段操纵大众舆论。其实施路径包括使用同步式或独立式机器人账号、依托单一社交媒体管理工具运营的多账号,或是对搜索引擎与社交网络算法进行各类干预,最终目的均为推广自身意识形态。尽管计算宣传正深刻影响现代社会,但其检测与量化工作却始终困难重重。此外,近年来大语言模型(Large Language Model)呈指数级增长,加之公众对信息过载的担忧与日俱增,令另类真相圈层的噪音环境愈发严峻,计算宣传的复杂度与规模亦将随之攀升,进一步加剧了其检测难度。 社交网络中的宣传内容常伪装成真实用户发布的正规新闻,巧妙地将真实用户与虚假账号融为一体。本研究旨在依托修辞学的核心宏观特征之一——重复特性,识别社交网络中操纵信息传播的行为。本次研究共纳入16个数据集,总容量达13GB:其中10个数据集围绕政治主题,剩余6个聚焦非政治主题(即大规模灾害场景),各数据集的推文数量从数万到数百万不等。研究通过对比两类数据集,挖掘出可区分这两类信息级联的统计特征与网络结构特征。这些特征同时涵盖话题标签(Hashtag)的重复分布、用户提及行为模式,以及网络拓扑结构。综合这些特征,我们得以在统计学意义上(p值=0.0001)有效区分两类信息级联。 除构建关联词汇与推文的信息级联二分图(bipartite graph)外,本研究还提出了一种量化指标,并验证了其可用于区分政治与非政治讨论场景。本方法不受信息级联的起源国家、语言或文化背景限制,仅依托重复特性的统计特征与推文中的词汇-二分图网络结构实现检测。



