遇见数据集

twitter Dataset

收藏
Zenodo2024-08-01 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The WE1S <code>twitter</code> dataset contains 5,024,756 tweets posted to Twitter between December 6th, 2013 and June 30th, 2019. The dataset is divided into subcollections based on the query terms "humanities", "liberal arts", "stem", "science", and "science-es" (that is a query for the presence of either "science" or "sciences"). Subcollections can be identified in the dataset from the value of the <code>metapath</code> property. The number of tweets in each subcollections is as follows: humanities: 1,705,038 liberal-arts: 7,663 stem: 865,156 science: 2,089,985 science-es: 356,914 The tweets are distributed over the following date range: 2013: 16,335 2014: 862,746 2015: 1,711,823 2016: 947,561 2017: 976,971 2018: 3,24,133 2019: 185,187 Collectively, the tweets represent the work of 1,886,739 distinct usernames. Each tweet's mentions, hashtags, and links are recorded, as well the number of likes and retweets. Unlike most other WE1S datasets, the Twitter dataset does not contain extracted features. Instead, it contains the original text of the tweet (the value of the <code>content</code> property, along with a <code>tidy_tweet</code> property, which contains the text of the tweet after preprocessing. Tweets were preprocessed using a modified form of the WE1S preprocessing algorithm. Details can be found in the WE1S Tweet-Suite repository. <em>(See WE1S Research Materials Overview for the relation between the project's "datasets" and "collections.")</em>

WE1S <code>twitter</code> 数据集收录了2013年12月6日至2019年6月30日期间发布至Twitter的5024756条推文。 该数据集根据检索词「humanities」(人文学科)、「liberal arts」(博雅学科)、「stem」(科学、技术、工程与数学,即STEM)、「science」(自然科学)以及「science-es」(即检索包含「science」或「sciences」的内容)划分为若干子数据集。可通过数据集中的<code>metapath</code>属性的取值识别各子数据集。 各子数据集的推文数量如下:humanities:1705038条;liberal-arts:7663条;stem:865156条;science:2089985条;science-es:356914条。 推文按年份分布如下:2013年:16335条;2014年:862746条;2015年:1711823条;2016年:947561条;2017年:976971条;2018年:3,24,133条;2019年:185187条。 所有推文共计涉及1886739个唯一用户名。每条推文的提及对象、话题标签(hashtags)与链接均被记录,同时包含点赞数与转发数。 与多数其他WE1S数据集不同,本Twitter数据集未包含提取后的特征,而是收录了推文的原始文本(即<code>content</code>属性的取值),同时附带<code>tidy_tweet</code>属性,该属性存储经过预处理后的推文文本。 推文预处理采用了经过改进的WE1S预处理算法,相关细节可于WE1S Tweet-Suite代码仓库中查阅。<em>(详见《WE1S研究材料概览》,了解该项目中「数据集」与「数据集集合」的对应关系。)</em>

提供机构:
Zenodo
创建时间:
2021-07-03
二维码
社区交流群
二维码
科研交流群
商业服务