[Tweets] 2022 Brazilian Presidential Elections
收藏资源简介:
2022 Brazilian Presidential Election This dataset contains 7,015,186 tweets from 951,602 users, extracted using 91 search terms over 36 days between August 1st and December 31st, 2022. All tweets in this dataset are in Brazilian Portuguese. Data Usage The dataset contains textual data from tweets, making it suitable for various NLP analyses, such as sentiment analysis, bias or stance detection, and toxic language detection. Additionally, users and tweets can be linked to create social graphs, enabling Social Network Analysis (SNA) to study polarization, communities, and other social dynamics. Extraction Method This data set was extracted using Twitter's (now X) official API—when Academic Research API access was still available—following the pipeline: 1. Twitter/X daily monitoring: The dataset author monitored daily political events appearing in Brazil's Trending Topics. Twitter/X has an automated system for classifying trending terms. When a term was identified as political, it was stored along with its date for later use as a search query. 2. Tweet collection using saved search terms: Once terms and their corresponding dates were recorded, tweets were extracted from 12:00 AM to 11:59 PM on the day the term entered the Trending Topics. A language filter was applied to select only tweets in Portuguese. The extraction was performed using the official Twitter/X API. 3. Data storage: The extracted data was organized by day and search term. If the same search term appeared in Trending Topics on consecutive days, a separate file was stored for each respective day. Further Information For more details, visit: - The repository- Dataset short paper: --- DOI: 10.5281/zenodo.14834669
2022年巴西总统选举数据集 本数据集包含来自951602名用户的7015186条推文,于2022年8月1日至12月31日期间,通过91个搜索词在36天周期内采集完成。 本数据集内所有推文均采用巴西葡萄牙语撰写。 数据用途 本数据集包含推文文本数据,适用于各类自然语言处理(Natural Language Processing, NLP)分析任务,例如情感分析、偏见与立场检测以及恶意语言检测。此外,可通过关联用户与推文构建社交图谱,借助社交网络分析(Social Network Analysis, SNA)研究舆论极化、社群结构及其他社会动态。 采集方法 本数据集通过Twitter(现更名为X)的官方应用程序编程接口(Application Programming Interface, API)采集,彼时学术研究API权限仍可申请使用,具体采集流程如下: 1. 推特/X日常监测:数据集作者每日监测巴西热搜榜(Trending Topics)中的政治事件。推特/X配备自动分类热搜词的系统,当识别到某一词汇为政治相关词汇时,将其及其出现日期存储下来,用作后续搜索查询词。 2. 基于存储搜索词的推文采集:记录下搜索词及其对应日期后,在该词汇登上热搜榜当日的00:00至23:59期间采集推文。采集过程中通过语言过滤器仅保留葡萄牙语推文,且全程使用推特/X官方API完成。 3. 数据存储:采集得到的数据按日期与搜索词进行分类存储。若同一搜索词连续多日登上热搜榜,则为每日分别生成独立存储文件。 更多信息 如需获取更多详情,请访问: - 数据集仓库与简短论文: --- DOI: 10.5281/zenodo.14834669



