twitter dataset
收藏资源简介:
该数据集是为名为User Modeling On Microblogging Websites的博士论文工作收集的,使用Twitter Streaming API在2015年11月4日至2016年1月12日期间收集了用户的实时推文。数据集包含177K用户和37M推文,用于研究识别Twitter上的主题权威。每个推文被分配零个、一个或多个主题。用户和Twitter ID已匿名化以遵守Twitter隐私政策,推文文本也被移除。数据集分为用户、推文和网络三个集合。
This dataset was collected for the doctoral thesis titled 'User Modeling On Microblogging Websites'. It utilized the Twitter Streaming API to gather real-time tweets from users between November 4, 2015, and January 12, 2016. The dataset comprises 177K users and 37M tweets, aimed at researching the identification of topic authorities on Twitter. Each tweet is assigned zero, one, or multiple topics. User and Twitter IDs have been anonymized in compliance with Twitter's privacy policy, and the tweet texts have been removed. The dataset is divided into three collections: users, tweets, and networks.
数据集概述
数据集名称
- twitter dataset
数据收集目的
- 用于名为“User Modeling On Microblogging Websites”的博士论文研究。
数据收集时间
- 2015年11月4日至2016年1月12日。
数据集规模
- 包含177K用户和37M tweets。
数据用途
- 用于研究Twitter上的主题权威识别。
数据内容
- 每个tweet可能分配有零个、一个或多个主题。
- 数据从MongoDB中导出。
- 用户和Twitter IDs已匿名处理以遵守Twitter隐私政策。
- 由于隐私原因,tweet文本已被移除。
数据集结构
- 包含三个集合:users, tweets, 和 network。
- tweets集合被分割成块,需要使用cat命令合并后才能恢复。
数据集样本
- 用户样本:包含_id, statusCount, friendsCount, followersCount, 和 tweet_count。
- tweet样本:包含_id, date, retweetCount, favCount, isRetweet, reTweetedTweetId, reTweetedUserId, hashtags, urls, userid, 和 topics。
- 网络样本:包含_id 和 followers。
引用要求
- 使用此数据集的研究应引用以下论文:
- Alp, Z. Z., & Öğüdücü, Ş. G. (2018). Identifying topical influencers on twitter based on user behavior and network topology. Knowledge-Based Systems, 141, 211-221.
- Alp, Z. Z., & Öğüdücü, Ş. G. (2019). Influence Factorization for identifying authorities in Twitter. Knowledge-Based Systems, 163, 944-954.




