遇见数据集

Datasets from `Discovering and analysing lexical variation in social media text'

收藏
Zenodo2020-07-30 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This repository contains the datasets that were used in the following three papers, which are also included within P. Shoemark's PhD dissertation `Discovering and analysing lexical variation in social media text': P. Shoemark, D. Sur, L. Shrimpton, I. Murray, and S. Goldwater. <em>Aye or naw, whit dae ye hink? Scottish independence and linguistic identity on social media.</em> 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL). 2017. P. Shoemark, J. Kirby, and S. Goldwater. <em>Topic and audience effects on distinctively Scottish vocabulary usage in Twitter data.</em> Workshop on Stylistic Variation at EMNLP. 2017. P. Shoemark, J. Kirby, and S. Goldwater. <em>Inducing a lexicon of sociolinguistic variables from code-mixed text.</em> Workshop on Noisy User Generated Text at EMNLP. 2018. Datasets consist of tab-separated-values files, in which rows correspond to tweets, with columns for user ID, tweet ID, and timestamp. The text of the tweets (and associated metadata) can be re-downloaded (in batches of 100 per request) using Twitter's GET Statuses/Lookup API endpoint <em>(NB: Tweets which have been deleted or made private since the original datasets were collected can <strong>not </strong>be re-downloaded, so it may not be possible to reconstruct the original datasets in their entirety). </em> Most of these datasets were originally drawn from the Sample endpoint of Twitter’s Streaming API (a.k.a. the ‘Spritzer’), which provides a random 1% sample of all public tweets in near real-time: <strong>GU Dataset: Geotagged-UK.zip</strong> <em>Tweets from Sept 2013 - Sept 2014 which are geotagged to locations within the UK.</em> The file <strong>GU_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by langid.py, are not retweets or quotes, and are geotagged to locations within the UK.<strong> </strong><em><strong>• Tweets: </strong>1,768,334<strong> • Unique Users: </strong>455,075 <strong>• </strong></em> The file <strong>GU.tsv </strong>contains the IDs for tweets in the <strong>final</strong> GU dataset that was used for the analyses in our EACL 2017 paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers. <em><strong>• Tweets: </strong>1,654,204<strong> • Unique Users: </strong>446,510 <strong>• </strong><sub>(the number of users in the GU dataset was slightly over-counted when reported in the paper; this is the actual number)</sub></em> <strong>GS Dataset: Geotagged-Scotland.zip </strong> <em>The subset of Tweets in the GU dataset which are geo-tagged to locations within Scotland, specifically.</em> The file <strong>GS_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by langid.py, are not retweets or quotes, and are geotagged to locations within Scotland. <em><strong>• Tweets: </strong>178,401<strong> • Unique Users: </strong>41,685 <strong>• </strong></em> The file <strong>GS.tsv </strong>contains the IDs for tweets in the <strong>final</strong> GS dataset that was used for the analyses in our EACL 2017 paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers. <em><strong>• Tweets: </strong>166,992<strong> • Unique Users: </strong>40,837 <strong>• </strong><sub>(the number of users in the GS dataset was slightly over-counted when reported in the paper; this is the actual number)</sub></em> <strong>IT Dataset &amp; Controls: Indyref-Tweets.zip</strong> <em>Tweets from Sept 2013 - Sept 2014 which contain hashtags relating to the 2014 Scottish Independence Referendum (plus 'control' tweets which are by the same users but do not contain referendum-related hashtags)</em> The file <strong>IT_</strong><strong>pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by langid.py, are not retweets or quotes, and contain at least one of 47 hashtags we identified as relating to the 2014 Scottish Independence Referendum (see paper for hashtag list). <em><strong>• Tweets: </strong>77,708<strong> • Unique Users: </strong>26,019 <strong>• </strong></em> The file <strong>IT</strong><strong>.tsv </strong>contains the IDs for tweets in the <strong>final</strong> IT dataset that was used for the analyses in our EACL 2017 paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers, and tweets which do not contain hashtags that we judged to <em>unambiguously</em> relate to the referendum. <em><strong>• Tweets: </strong>59,664 <strong>• Unique Users: </strong>18,589 <strong>• </strong></em> The file <strong>IT_controls_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by langid.py, are not retweets or quotes, and do <em><strong>not</strong></em> contain any of the hashtags we identified as relating to the 2014 Scottish Independence Referendum, but are authored by a user who has <em><strong>also</strong> </em>authored a tweet in <strong>IT_</strong><strong>pre-filtering.tsv</strong>. <em><strong>• Tweets: </strong>1,354,701 <strong>• Unique Users: </strong>26,019 <strong>• </strong></em> The file <strong>IT_controls</strong><strong>.tsv </strong>contains the IDs for tweets in the <strong>final</strong> set of Control tweets that was used for the analyses in our EACL 2017 paper, i.e. tweets which do not contain referendum-related hashtags but are authored by users who also have also authored tweets in <strong>IT</strong><strong>.tsv</strong>. <em><strong>• Tweets: </strong>881,679 <strong>• Unique Users: </strong>18,589 <strong>• </strong></em> <strong>SG-Users’ and IH-Users' Autumn 2014 Timeline Datasets: Autumn-2014_Timelines.zip</strong> <em>Complete tweet histories from Aug-Oct 2014 for users from the GS and IT datasets.</em> The file <strong>SG-Users_Autumn_2014_timelines_pre-filtering.tsv </strong>contains the IDs for tweets which were posted in August, September, or October 2014 by users from the GS dataset, i.e. users we know to have used Scottish geotags. This dataset is not restricted to tweets which appear in the ‘Spritzer’ sample; instead it consists of complete User Timelines for the months concerned, retrieved using the statuses/user timeline endpoint of Twitter’s REST API in March 2017. Because there are limits on the number of tweets that can be retrieved using this endpoint, we were not able to retrieve complete Autumn 2014 tweet histories for <em>all</em> of the users in the GS dataset. <em><strong>• Tweets: </strong>3,014,029 </em> <em><strong>• Unique Users: </strong>18,274 <strong>• </strong></em> The file <strong>SG-Users_Autumn_2014_timelines.tsv </strong>contains the IDs for tweets in the <strong>final</strong> SG-Users dataset that was used for the analyses in our StyleVar 2017 paper. This dataset consists only of tweets which contain at least one instance of one of 50 lexical variables that were the focus of the study, and has also undergone various other filtering steps; see the paper for full details. <em><strong>• Tweets: </strong>1,112,931</em> <em><strong>• Unique Users: </strong>10,103 <strong>• </strong></em> The file <strong>IH-Users_Autumn_2014_timelines_pre-filtering.tsv </strong>contains the IDs for tweets which were posted in August, September, or October 2014 by users from the IT dataset, i.e. users we know to have used Indyref-related hashtags. This dataset was collected in the same manner as SG-Users_Autumn_2014_timelines_pre-filtering.tsv; however, due to an error in this process, <strong>the IDs of most of the tweets in this dataset were not recorded</strong>. For such tweets the tweet ID column instead contains a placeholder tweet ID of the form _&lt;user_ID&gt;_&lt;month&gt;_&lt;integer&gt;, where the integer denotes the tweet's position in the reverse-chronological list of tweets that were retrieved for that user from that month (e.g. _147527441_09_286 is the placeholder tweet ID we assigned to the 286th September tweet we retrieved from the user whose Account ID is 147527441). Unfortunately, therefore, the tweets in this file whose 'IDs' begin with an underscore cannot be straightforwardly re-downloaded using Twitter's free GET Statuses/Lookup API endpoint; but since their user IDs and timestamps are intact, it would still be possible to retrieve them using the (paid-for) Historical APIs. <em><strong>• Tweets: </strong>6,997,858 <strong>• Tweets whose IDs were recorded: </strong>288,394</em><em><strong> </strong> <strong>• Unique Users: </strong>14,645</em><em> <strong>• </strong></em> The file <strong>IH-Users_Autumn_2014_timelines.tsv </strong>contains the IDs for tweets in the <strong>final</strong> IH-Users dataset that was used for the analyses in our StyleVar 2017 paper. Like the SG-Users dataset, this dataset consists only of tweets which contain at least one instance of one of 50 lexical variables that were the focus of the study, and has also undergone various other filtering steps; see the paper for full details. As with the pre-filtered version, most of the tweet IDs are unfortunately missing in this dataset. <em><strong>• Tweets: </strong>2,165,320 <strong>• Tweets whose IDs were recorded: </strong>115,366</em><em><strong> </strong> <strong>• Unique Users: </strong>10,784 <strong>• </strong></em> <strong>US Geotags: Geotagged-USA.zip</strong> <em>Tweets from June 2013 - July 2016 which are geotagged to locations within the USA.</em> The file <strong>GUSA.tsv</strong> contains all tweets from the ‘Spritzer’ sample which were posted between June 30th 2013 to July 1st 2016, are classified as English by langid.py, are not retweets, do not contain urls or embedded media, are not by users with more than 1000 friends or followers, and are geotagged to locations within the USA. This dataset (along with the GU Dataset) was used in our WNUT 2018 paper. <em><strong>• Tweets: </strong></em> <em>8,375,573 </em> <em><strong>• Unique Users: </strong>1</em>,<em>826,260</em><em> <strong>• </strong></em>

本仓库收录了三篇学术论文所使用的数据集,这些数据集同时被纳入P. Shoemark的博士学位论文《社交媒体文本中词汇变异的发现与分析》(*Discovering and analysing lexical variation in social media text*)。三篇论文分别为: 1. P. Shoemark、D. Sur、L. Shrimpton、I. Murray与S. Goldwater。*《Aye or naw, whit dae ye hink? Scottish independence and linguistic identity on social media》*,发表于第15届欧洲计算语言学协会(Association for Computational Linguistics, EACL)大会,2017年。 2. P. Shoemark、J. Kirby与S. Goldwater。*《Topic and audience effects on distinctively Scottish vocabulary usage in Twitter data》*,发表于EMNLP文体变异研讨会(Workshop on Stylistic Variation),2017年。 3. P. Shoemark、J. Kirby与S. Goldwater。*《Inducing a lexicon of sociolinguistic variables from code-mixed text》*,发表于EMNLP噪声用户生成文本研讨会(Workshop on Noisy User Generated Text),2018年。 本数据集均采用制表符分隔值(Tab-Separated Values, TSV)格式存储,每行对应一条推文,字段依次为用户ID、推文ID与时间戳。推文文本及相关元数据可通过Twitter的GET Statuses/Lookup应用程序编程接口(API)批量重新获取(单次请求最多支持100条推文)。**注意**:自原始数据集采集之日起已删除或设为私密的推文无法重新获取,因此可能无法完整重建原始数据集。 本批次数据集绝大多数最初源自Twitter流式API(Streaming API)的Sample端点,俗称“Spritzer”,该接口可近乎实时地随机抽取全部公开推文的1%样本: ### **GU数据集:Geotagged-UK.zip** 该数据集为2013年9月至2014年9月期间、地理标签位于英国境内的推文。 文件`GU_pre-filtering.tsv`包含所有来自“Spritzer”流的推文ID,这些推文发布于2013年9月1日至2014年9月30日之间,经langid.py识别为英语,非转发或引用推文,且地理标签位于英国境内。 - 推文总量:1,768,334 - 唯一用户数:455,075 文件`GU.tsv`包含经额外启发式预处理过滤(剔除机器人与垃圾发送者发布的推文)后,用于EACL 2017论文分析的**最终版**GU数据集的推文ID。 - 推文总量:1,654,204 - 唯一用户数:446,510 *(论文中报告的GU数据集用户数存在轻微高估,此为实际数值)* ### **GS数据集:Geotagged-Scotland.zip** 该数据集为GU数据集中地理标签位于苏格兰境内的推文子集。 文件`GS_pre-filtering.tsv`包含所有来自“Spritzer”流的推文ID,这些推文发布于2013年9月1日至2014年9月30日之间,经langid.py识别为英语,非转发或引用推文,且地理标签位于苏格兰境内。 - 推文总量:178,401 - 唯一用户数:41,685 文件`GS.tsv`包含经额外启发式预处理过滤(剔除机器人与垃圾发送者发布的推文)后,用于EACL 2017论文分析的**最终版**GS数据集的推文ID。 - 推文总量:166,992 - 唯一用户数:40,837 *(论文中报告的GS数据集用户数存在轻微高估,此为实际数值)* ### **IT数据集与对照组:Indyref-Tweets.zip** 该数据集为2013年9月至2014年9月期间、包含与2014年苏格兰独立公投相关话题标签(hashtag)的推文,以及来自相同用户但未包含公投相关话题标签的“对照”推文。 文件`IT_pre-filtering.tsv`包含所有来自“Spritzer”流的推文ID,这些推文发布于2013年9月1日至2014年9月30日之间,经langid.py识别为英语,非转发或引用推文,且包含我们识别出的47个与2014年苏格兰独立公投相关的话题标签之一(话题标签列表详见论文)。 - 推文总量:77,708 - 唯一用户数:26,019 文件`IT.tsv`包含经额外启发式预处理过滤(剔除机器人与垃圾发送者发布的推文,以及未包含我们判定为**明确**与公投相关的话题标签的推文)后,用于EACL 2017论文分析的**最终版**IT数据集的推文ID。 - 推文总量:59,664 - 唯一用户数:18,589 文件`IT_controls_pre-filtering.tsv`包含所有来自“Spritzer”流的推文ID,这些推文发布于2013年9月1日至2014年9月30日之间,经langid.py识别为英语,非转发或引用推文,**未**包含任何与2014年苏格兰独立公投相关的话题标签,但发布者同时在`IT_pre-filtering.tsv`中发布过推文。 - 推文总量:1,354,701 - 唯一用户数:26,019 文件`IT_controls.tsv`包含经处理后,用于EACL 2017论文分析的**最终版**对照推文数据集的推文ID,即未包含公投相关话题标签,但发布者同时在`IT.tsv`中发布过推文的用户所发布的推文。 - 推文总量:881,679 - 唯一用户数:18,589 ### **SG用户与IH用户2014年秋季时间线数据集:Autumn-2014_Timelines.zip** 该数据集包含GS与IT数据集用户在2014年8月至10月期间的完整推文历史。 文件`SG-Users_Autumn_2014_timelines_pre-filtering.tsv`包含GS数据集用户(即已知使用苏格兰地理标签的用户)在2014年8月、9月或10月发布的推文ID。该数据集并非仅限于“Spritzer”样本中的推文,而是通过Twitter REST API的statuses/user timeline接口于2017年3月获取的对应月份的完整用户时间线。由于该接口的推文获取数量限制,我们无法获取GS数据集中**所有**用户的2014年秋季完整推文历史。 - 推文总量:3,014,029 - 唯一用户数:18,274 文件`SG-Users_Autumn_2014_timelines.tsv`包含用于StyleVar 2017论文分析的**最终版**SG用户数据集的推文ID。该数据集仅包含至少包含50个本研究关注的词汇变量之一的推文,并经过了多种其他过滤步骤;完整细节详见论文。 - 推文总量:1,112,931 - 唯一用户数:10,103 文件`IH-Users_Autumn_2014_timelines_pre-filtering.tsv`包含IT数据集用户(即已知使用与独立公投相关话题标签的用户)在2014年8月、9月或10月发布的推文ID。该数据集的收集方式与`SG-Users_Autumn_2014_timelines_pre-filtering.tsv`一致;但由于采集过程中的错误,该数据集中**大部分推文的ID未被记录**。对于此类推文,推文ID列将使用占位符,格式为`_<user_ID>_<month>_<integer>`,其中整数表示从该用户该月获取的推文按逆序排列后的位置(例如`_147527441_09_286`表示我们从账号ID为147527441的用户处获取的第286条9月推文的占位符ID)。因此,此类以下划线开头的“ID”无法通过Twitter免费的GET Statuses/Lookup API接口直接重新获取;但由于其用户ID与时间戳信息完整,仍可通过付费的历史数据接口获取推文。 - 推文总量:6,997,858 - 已记录ID的推文:288,394 - 唯一用户数:14,645 文件`IH-Users_Autumn_2014_timelines.tsv`包含用于StyleVar 2017论文分析的**最终版**IH用户数据集的推文ID。与SG用户数据集类似,该数据集仅包含至少包含50个本研究关注的词汇变量之一的推文,并经过了多种其他过滤步骤;完整细节详见论文。与预处理版本相同,该数据集中的大部分推文ID缺失。 - 推文总量:2,165,320 - 已记录ID的推文:115,366 - 唯一用户数:10,784 ### **美国地理标签数据集:Geotagged-USA.zip** 该数据集为2013年6月至2016年7月期间、地理标签位于美国境内的推文。 文件`GUSA.tsv`包含所有来自“Spritzer”样本的推文,这些推文发布于2013年6月30日至2016年7月1日之间,经langid.py识别为英语,非转发推文,不包含链接或内嵌媒体,发布者的好友与粉丝数均不超过1000,且地理标签位于美国境内。该数据集(与GU数据集一同)被用于我们的WNUT 2018论文。 - 推文总量:8,375,573 - 唯一用户数:1,826,260

提供机构:
Zenodo
创建时间:
2019-11-26
二维码
社区交流群
二维码
科研交流群
商业服务