遇见数据集

Arabic Tweets on the Ghada Aoun Judicial Case: Manual Emotional/Rational Annotations and Full Corpus (Lebanon, April–May 2021)

收藏
Zenodo2026-05-31 更新2026-06-05 收录
官方服务:

资源简介:

This dataset accompanies the study Artificial Intelligence for Social Media Analysis: Detecting Pseudo-Facts in Digital Environments. It contains Arabic-language tweets (predominantly Lebanese dialect, with Modern Standard Arabic) collected from X (formerly Twitter) via the Twitter API between 20 April and 6 May 2021, during the public controversy surrounding Judge Ghada Aoun in Lebanon. The data support reproduction of the supervised text-classification experiments reported in the article, which distinguish emotional from rational discourse as an indicator of how pseudo-factual claims are framed and circulated. The deposit comprises two files. The first, a manually labelled sample of 600 tweets (443 emotional, 157 rational), provides the training and evaluation set; it includes an intermediate three-way code (emotional, political, judicial), the coder's notes, and the final binary label. The second is the full corpus of 26,489 tweets to which the trained model was applied. The full annotation protocol — operational definitions, decision rules, and worked examples — is given in Section 3.1 of the associated article. All tweet text has been anonymised: account handles were replaced with "@user" and URLs with "[URL]", and the files contain no usernames, user identifiers, display names, profile images, geolocation, or timestamps. Hashtags and the names of public figures central to the event are retained. A README/data dictionary documents the variables, coding scheme, licence, and known limitations (notably that the record identifiers are sequential indices rather than Twitter/X status IDs). The data are released under CC BY 4.0; the underlying content remains subject to X's Terms of Service.

提供机构:
Zenodo
创建时间:
2026-05-31
二维码
社区交流群
二维码
科研交流群
商业服务