Arabic Tweets on the Ghada Aoun Judicial Case: Manual Emotional/Rational Annotations and Full Corpus (Lebanon, April–May 2021)
收藏资源简介:
This dataset accompanies the study Artificial Intelligence for Social Media Analysis: Detecting Pseudo-Facts in Digital Environments. It contains Arabic-language tweets (predominantly Lebanese dialect, with Modern Standard Arabic) collected from X (formerly Twitter) via the Twitter API between 20 April and 6 May 2021, during the public controversy surrounding Judge Ghada Aoun in Lebanon. The data support reproduction of the supervised text-classification experiments reported in the article, which distinguish emotional from rational discourse as an indicator of how pseudo-factual claims are framed and circulated. The deposit comprises two files. The first, a manually labelled sample of 600 tweets (443 emotional, 157 rational), provides the training and evaluation set; it includes an intermediate three-way code (emotional, political, judicial), the coder's notes, and the final binary label. The second is the full corpus of 26,489 tweets to which the trained model was applied. The full annotation protocol — operational definitions, decision rules, and worked examples — is given in Section 3.1 of the associated article. All tweet text has been anonymised: account handles were replaced with "@user" and URLs with "[URL]", and the files contain no usernames, user identifiers, display names, profile images, geolocation, or timestamps. Hashtags and the names of public figures central to the event are retained. A README/data dictionary documents the variables, coding scheme, licence, and known limitations (notably that the record identifiers are sequential indices rather than Twitter/X status IDs). The data are released under CC BY 4.0; the underlying content remains subject to X's Terms of Service.
本数据集配套于研究论文《用于社交媒体分析的人工智能:检测数字环境中的伪事实》。数据集包含2021年4月20日至5月6日期间,通过Twitter API从X(原Twitter)平台采集的阿拉伯语推文,其中以黎巴嫩方言为主,兼含现代标准阿拉伯语,采集时段正值黎巴嫩围绕法官加达·乌恩(Ghada Aoun)爆发的公共争议事件。该数据集可复现论文中报道的监督式文本分类实验,该实验以区分情感话语与理性话语为指标,用以分析伪事实主张的构建与传播路径。 本数据集包含两个文件:其一为经人工标注的600条推文样本(其中443条为情感类,157条为理性类),作为训练与评估集;该样本包含中间三元编码(情感、政治、司法)、标注者笔记以及最终的二元标签。其二为经训练模型推演应用的全部推文语料库,共计26489条。完整的标注规范——包括操作定义、决策规则与标注示例——详见相关论文的3.1章节。 所有推文文本均已完成匿名化处理:账号用户名替换为"@user",URL链接替换为"[URL]",文件中未包含用户名、用户标识符、显示名称、头像、地理位置或时间戳。与该事件核心相关的话题标签与公众人物姓名予以保留。本数据集附带README文件(数据字典),用于说明变量、编码方案、授权协议以及已知局限性(具体而言,记录标识符为连续索引,而非Twitter/X的推文状态ID)。本数据集采用CC BY 4.0协议发布,其底层内容仍受X的服务条款约束。



