End of Term 2016 U.S. Government Twitter Archive
收藏资源简介:
This dataset contains the tweet ids of 5,655,632 tweets that were collected from approximately 3000 Twitter accounts affiliated with the U.S. government. They were collected between October 21, 2016 and January 21, 2017 from the Twitter API using Social Feed Manager. This dataset was created as part of the End of Term Web Archiving initiative. The lists of accounts came from the U.S. Digital Registry and by public submissions. These tweets were collected using the POST statuses/user_timeline method of the Twitter REST API. There is a README.txt file containing additional documentation on how it was collected. The GET statuses/lookup method supports retrieving the complete tweet for a tweet id (known as hydrating). Tools such as Twarc or Hydrator can be used to hydrate tweets. Per Twitter’s Developer Policy, tweet ids may be publicly shared; tweets may not. Questions about this dataset can be sent to sfm@gwu.edu. George Washington University researchers should contact us for access to the tweets. This work is supported by grant #NARDI-14-50017-14 from the National Historical Publications and Records Commission.
本数据集包含5,655,632条推文的ID,其采集自约3000个隶属于美国联邦政府的Twitter账号。本次数据采集于2016年10月21日至2017年1月21日期间完成,通过Social Feed Manager工具调用Twitter API接口实现。本数据集为「任期结束网络归档(End of Term Web Archiving)」倡议的组成部分,所涉账号列表源自美国数字登记处(U.S. Digital Registry)及公众提交的账号信息。本次推文采集采用了Twitter REST API的POST statuses/user_timeline接口方法。数据集附带README.txt文件,其中包含数据采集流程的补充说明文档。GET statuses/lookup接口支持通过推文ID检索完整推文内容,该操作被称为「推文补全(hydrating)」,可使用Twarc或Hydrator等工具完成该补全流程。根据Twitter开发者政策,推文ID可公开共享,但完整推文内容不得对外分享。若对本数据集存在疑问,请发送邮件至sfm@gwu.edu。乔治华盛顿大学(George Washington University)的研究人员如需获取完整推文内容,请与我方联系。本研究工作由国家历史出版物与档案委员会(National Historical Publications and Records Commission)提供的编号为NARDI-14-50017-14的资助项目支持。



