遇见数据集

Annotated Dataset of History-related Tweets

收藏
Zenodo2021-09-19 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This repository contains tweet IDs and their 5 types of contextual information including 1) hashtags, 2) their categories, 3) entities obtained by NERD, 4) time-references normalized by Heideltime, and 5) Web categories for URLs attached with history-related hashtag that are related to history and that were collected for the purpose of analyzing how history-related content is disseminated in online social networks. Our IJDL paper shows the analysis results. The preliminary version of the analysis report is available here. We used the Twitter official search API provided by Twitter to collect tweets. Note that three kinds of tweets are typically found in Twitter: tweets, retweets and quote tweets. Tweet is an original text issued as a post by a Twitter user. A retweet is a copy of an original tweet for the purpose of propagating the tweet content to more users (i.e., one's followers). Finally, a quote tweet copies the content of another tweet and allows also to add new content. A quote tweet is sometimes called a retweet with a comment. In this work, we simply treat all quote tweets as original tweets since they include additional information/text. There were however only 1,877 (0.2%) tweets recognized as quote tweets in our dataset. To collect tweets that refer to the past or are related to collective memory of past events/entities, we performed hashtag based crawling together with bootstrapping procedure. <br> At the beginning, we gathered several historical hashtags selected by experts (e.g. <strong>#HistoryTeacher</strong>, <strong>#history</strong>, <strong>#WmnHist</strong>). <br> In addition, we prepared several hashtags that are commonly used when referring to the past: <strong>#onthisday</strong>, <strong>#thisdayinhistory</strong>, <strong>#throwbackthursday</strong>, <strong>#otd</strong>. We then collected tweets that contain these hashtags by using Twitter official search API. The collected tweets were issued from 8 March 2016 to 2 July 2018. <br> Bootstrapping allowed us to search for other hashtags frequently used with the seed hashtags. The tweets tagged by such hashtags were then included into the seed set after the manual inspection of all the discovered hashtags as of their relation to the history, and filtering ones that are unrelated. <br> In total, we gathered 147 history-related hashtags which allowed us to collect 2,370,252 tweet IDs pointing to 882,977 tweets and 1,487,275 re-tweets. Related papers: Yasunobu Sumikawa, Adam Jatowt, and Marten During, <strong>"Digital History meets Microblogging: Analyzing Collective Memories in Twitter"</strong>, In Proceedings of the 18th ACM/IEEE-CS Joint Conference on Digital Libraries, JCDL'18, IEEE/ACM, pp. 213 -- 222, 2018. [paper] Yasunobu Sumikawa and Adam Jatowt, <strong>"Analyzing History-related Posts in Twitter"</strong>, International Journal on Digital Libraries, Springer, 2020. https://doi.org/10.1007/s00799-020-00296-2 [paper][dataset] Yasunobu Sumikawa and Adam Jatowt, <strong>"Annotated Dataset of History-related Tweets"</strong>, Data in Brief, Vol. 38, pp. 107344, Elsevier, 2021. [paper]

本仓库存储了推文ID及其五类上下文信息,分别为:1)话题标签(hashtags),2)话题标签分类,3)由NERD提取的实体,4)经Heideltime标准化后的时间参考信息,5)绑定历史相关话题标签的URL的网页分类。本数据集采集自在线社交网络,用于分析历史相关内容的传播规律。 我们发表于《国际数字图书馆期刊(International Journal on Digital Libraries,IJDL)》的论文展示了本数据集的分析结果,分析报告的预印版本可在此处获取。 我们通过Twitter官方提供的搜索API采集推文。需注意,Twitter平台上通常存在三类推文:原创推文(tweet)、转发推文(retweet)与引用推文(quote tweet)。 原创推文指Twitter用户发布的原生文本内容;转发推文是对原创推文的转发,用于将内容传播给更多用户(即发布者的关注者);引用推文则在转发另一推文的同时可添加新的文本内容,此类推文有时也被称为“带评论的转发推文”。 本研究中,由于引用推文包含额外的信息或文本,我们将所有引用推文均视为原创推文进行处理。但在本数据集中,仅有1,877(0.2%)条推文被标注为引用推文。 为采集涉及过去事件或与过往事件/实体的集体记忆相关的推文,我们采用了基于话题标签的爬取方法结合自举(bootstrapping)流程。 最初,我们征集了专家选定的若干历史类话题标签,例如**#HistoryTeacher**、**#history**、**#WmnHist**。 此外,我们还准备了若干常用于提及过往事件的话题标签:**#onthisday**、**#thisdayinhistory**、**#throwbackthursday**、**#otd**。随后我们通过Twitter官方搜索API采集了包含上述话题标签的推文,采集时间范围为2016年3月8日至2018年7月2日。 通过自举流程,我们可以挖掘出与种子话题标签高频关联的其他话题标签。在对所有新发现的话题标签进行人工审核,筛选出与历史相关的标签后,将其对应的推文纳入种子数据集。 最终我们共收集到147个历史相关话题标签,据此采集到2,370,252条推文ID,对应882,977条原创推文与1,487,275条转发推文。 相关论文: 1. Yasunobu Sumikawa、Adam Jatowt与Marten During,**《数字史学与微博:分析Twitter中的集体记忆》**,发表于第18届ACM/IEEE-CS联合数字图书馆会议(JCDL'18),IEEE/ACM,2018年,第213-222页。[论文] 2. Yasunobu Sumikawa与Adam Jatowt,**《分析Twitter中的历史相关帖子》**,《国际数字图书馆期刊(International Journal on Digital Libraries)》,Springer,2020年。https://doi.org/10.1007/s00799-020-00296-2 [论文][数据集] 3. Yasunobu Sumikawa与Adam Jatowt,**《历史相关推文标注数据集》**,《数据简报(Data in Brief)》,第38卷,第107344页,Elsevier,2021年。[论文]

提供机构:
Zenodo
创建时间:
2021-04-01
二维码
社区交流群
二维码
科研交流群
商业服务