U.S. Government Tweet Ids
收藏资源简介:
This dataset contains the tweet ids of 9,673,959 tweets for approximately 3400 U.S. government accounts. These are accounts that are associated with federal government agencies, not individuals. They were collected between January 20, 2017 and July 20, 2018 from the GET statuses/user_timeline method of the Twitter API using Social Feed Manager. There is a README.txt file containing additional documentation on how this dataset was collected. There is also an accounts.csv file listing the Twitter accounts that were collected. Note that accounts have been added and deleted during the collection period. This may contain tweets from accounts that were erroneously collected (e.g., accounts of government officials or accounts that were previously government accounts but since taken over by non-government users). The GET statuses/lookup method supports retrieving the complete tweet for a tweet id (known as hydrating). Tools such as Twarc or Hydrator can be used to hydrate tweets. Per Twitter’s Developer Policy, tweet ids may be publicly shared for academic purposes; tweets may not. We intend to update this dataset periodically. Questions about this dataset can be sent to sfm@gwu.edu. George Washington University researchers should contact us for access to the tweets.
本数据集涵盖约3400个美国政府官方账号对应的9,673,959条推文的推文ID(tweet IDs)。此类账号均与联邦政府机构相关联,而非个人账号。本数据集通过社交馈送管理器(Social Feed Manager)调用推特应用程序编程接口(Twitter API)的GET statuses/user_timeline方法,于2017年1月20日至2018年7月20日期间完成采集。数据集附带README.txt文件,其中包含有关本数据集采集流程的补充说明文档;同时附带accounts.csv文件,其中列出了本次采集的所有推特账号。请注意,在采集周期内,部分账号曾进行新增或删除操作。本数据集可能包含误采集账号的推文,例如政府官员账号,或此前为政府账号但后续被非政府用户接管的账号。GET statuses/lookup接口支持通过推文ID获取完整推文内容,该操作被称为hydrating。可使用Twarc或Hydrator等工具完成推文的hydrating操作。根据推特开发者政策,出于学术目的可公开分享推文ID,但不得公开完整推文内容。我们计划定期对本数据集进行更新。有关本数据集的相关疑问可发送至邮箱sfm@gwu.edu。乔治华盛顿大学的研究人员若需获取完整推文内容,请与我们联系。



