遇见数据集

Covid Twitter Emotion Analysis Data

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

Twitter data was collected using Twitter’s Application Programming Interface(API) and Tweepy, a python library to access the twitter API. Certain keywords related to COVID’19 like Coronavirus, ncov, Wuhan, China, Covid-19, Epidemic, Pandemic, SocialDistancing, etc. were used to collect the tweets. Only the tweets that were in English and the ones that had a geo-tag were collected. During the exploratory data analysis, we noticed that a number of tweets consisted of only certain words and not proper sentences and analyzing the emotion of such tweets might not give us a proper overview of the emotions. Thus, only the tweets with at least 6 words in them were used. This significantly reduced the number of tweets collected. Finally, we had over 1 million tweets over the span of February, March, April, May, and June. The tweets were then further processed to remove all the HTML text, ‘@’ mentions, URL links, and #hashtags. The data was analyzed using a machine learning model and tweets were categorized into various emotions. The dataset provides the count of tweets per country per emotion for 5 months.

本数据集的推特数据通过推特应用程序编程接口(Application Programming Interface,API)与Python库Tweepy获取。研究人员以与新型冠状病毒肺炎(COVID-19)相关的关键词,如冠状病毒(Coronavirus)、ncov、武汉、中国、Covid-19、流行病(Epidemic)、大流行病(Pandemic)、社交距离(SocialDistancing)等作为检索条件抓取推文,仅收录英语语种且带有地理标签(geo-tag)的内容。在探索性数据分析阶段,团队发现大量推文仅由单个单词构成而非完整语句,此类推文的情感分析难以反映真实的情绪全貌,因此仅保留包含至少6个单词的推文,该筛选条件大幅缩减了初始推文规模。最终在2月至6月的五个月周期内,共收集到超100万条推文。随后对推文进行进一步预处理,移除所有HTML文本、@提及内容、URL链接与话题标签(hashtags)。 本数据集采用机器学习模型对采集到的数据进行分析,将推文划分为多种情感类别。数据集提供了五个月内,每个国家对应各类情感的推文数量统计结果。

创建时间:
2020-07-16
二维码
社区交流群
二维码
科研交流群
商业服务