遇见数据集

Elite Twitter Polarization Dataset

收藏
Zenodo2026-02-17 更新2026-05-26 收录
官方服务:

资源简介:

README – Elite Twitter Polarization Dataset (2010–2021) Overview This dataset contains annotated Twitter activity from nine globally influential public figures spanning the period 1 January 2010 to 31 December 2021. Each record represents either an original tweet or a retweet authored by one of the selected individuals and includes sentiment metrics, polarization classification, and topic assignments. Five of the nine figures are among the most-followed users on the platform, each with over 100 million followers, ensuring coverage of both the highest-reach accounts and diverse professional domains. Sample The sample includes globally recognized individuals from politics, business, sports, music, and philanthropy. Selection criteria were: Global public recognition. High follower counts (ranging from 0.19 million to 200 million as of July 2024). Active presence on Twitter between 2010 and 2021. The nine individuals in this dataset are: No. Name ID Followers (Million, July 2024) Profession Tweets Retweets 1 Elon Musk @elonmusk 200 Businessman 2,309 1,687 2 Barack Obama @BarackObama 131.7 44th President of the US 14,032 1,886 3 Cristiano Ronaldo @Cristiano 112.1 Football Player 3,383 117 4 Katy Perry @katyperry 106.3 Musician 8,073 378 5 Narendra Modi @narendramodi 100 Prime Minister of India 26,807 1,458 6 Bill Gates @BillGates 66.3 Businessman 3,540 295 7 Melinda Gates @melindagates 2.4 Philanthropist 3,283 365 8 Marc Benioff @Benioff 1.1 Businessman 9,316 23,189 9 John Collison @collision 0.196 Entrepreneur 1,899 1,566 Folder Structure The dataset is organised by individual. Each individual has a dedicated folder containing two Excel files: /[Person_Name]/ Tweets.xlsx Retweets.xlsx Tweets.xlsx – All original tweets posted by the individual between 2010-01-01 and 2021-12-31. Retweets.xlsx – All retweets posted by the individual during the same period. File Structure Each Excel file contains the following 10 columns: Creation Date – Date of tweet/retweet creation. ID – Unique tweet/retweet ID issued by Twitter. Sentiment Negative Score – Negative sentiment score (VADER algorithm). Sentiment Positive Score – Positive sentiment score (VADER algorithm). Sentiment Compound Score – Compound sentiment score (VADER; range -1 to 1). Head Topic – Higher-level thematic topic (e.g., “Politics and Governance”, “Sports and Leisure”) generated via Grok 3 large language model. Topic – Surface-level topic extracted via GPT-3.5 (e.g., “Call to Action for Disaster Relief in Haiti”). Stance – Indicates whether the author expresses a clear opinion or position on an issue (stance-taking). Controversy – Indicates whether the topic is likely to provoke disagreement or debate (controversial). Is Polarized – Boolean classification based on the two criteria above: Yes (Polarized): Content meets both criteria (stance-taking and controversial). No (Non-Polarized): Content meets neither criterion. NLA (No Label Assigned): Content meets only one of the two criteria (excluded in order to maintain a clear distinction between polarized and non-polarized posts). Data Collection and Annotation Source: Collected using Twitter API v2 with Academic Research access. Period: 1 January 2010 – 31 December 2021. Preprocessing: Removal of URLs, emojis, and non-text characters prior to entity and topic extraction. Polarization Annotation: GPT-3.5-Turbo was applied with a standardized prompt to detect both controversy and stance-taking. Posts meeting both criteria were classified as Polarized. Posts meeting neither criterion were classified as Non-Polarized. Posts meeting only one criterion were excluded (NLA – No Label Assigned) to ensure a strict separation between categories. To validate accuracy, 558 posts were manually reviewed (with Grok 4 as aid). Agreement with GPT-3.5 was 94.4% for topic assignment and 99.5% for polarization status, confirming high reliability. Topic Modeling: Stage 1: GPT-3.5 extracted surface-level topics. Stage 2: Grok 3 clustered these into thematic “Head Topics.” Stage 3: Manual validation on 270 posts (with Grok 4 as aid) showed 77.8% agreement with Grok 3’s thematic clustering, indicating substantial alignment. Sentiment Analysis: Conducted using VADER (Valence Aware Dictionary and sEntiment Reasoner). Negative, Positive, and Compound scores calculated for each post. Ethics and Compliance This dataset complies with Twitter’s Developer Policy by including only tweet IDs and derived annotations, not full tweet text. Users must hydrate tweet IDs using the Twitter API to retrieve original content, subject to Twitter’s terms of service. This structure ensures ethical sharing and reproducibility.

提供机构:
Zenodo
创建时间:
2025-08-23
二维码
社区交流群
二维码
科研交流群
商业服务