Monthly Samples of German Tweets
收藏资源简介:
This dataset contains German tweets and Twitter accounts recorded from the public Twitter Streaming API using the following filters: terms: <em>'a'</em>, <em>'e'</em>, <em>'i'</em>, <em>'o'</em>, <em>'u'</em>, and <em>'n'</em> language: <em>'de'</em> This filter combination should record (almost) all German tweets (in German it is very unlikely that terms do not contain vowels or the frequently used character <em>'n'</em>). This dataset might be useful for the following use cases: Natural language processing (focussing on Twitter specifics in German, there exist only little German datasets) Social Network Analysis (Twitter network) Identifying behavioural patterns (retweeting, quoting, replying, hate speech, ...) Sharing political (or other domain-specific) content Bot detection and more ... This dataset will be updated monthly. Each sample (starting in April 2019) will follow the following naming pattern: german-tweet-sample-<em><YEAR></em>-<em><MONTH></em>.zip (size: ~ 1GB) It will contain several bunches of recorded JSON gzipped files. Each bunch of records contains approximately 150k recorded tweets/accounts (size: ~ 18MB).
本数据集收录通过公开Twitter流式API采集的德语推文与Twitter账号,采集时采用以下过滤规则:关键词为元音字母'a'、'e'、'i'、'o'、'u'及字母'n',语言过滤条件设为德语('de')。该过滤组合可覆盖(几乎)全部德语推文——在德语语境中,极少有词汇不包含元音或高频字符'n'。 本数据集可应用于如下场景: 1. 自然语言处理(聚焦德语平台Twitter的语言特性,目前德语相关数据集仍较为稀缺) 2. 社交网络分析(Twitter网络结构) 3. 行为模式识别(涵盖转发、引用、回复、仇恨言论识别等) 4. 政治(或其他垂直领域)内容传播分析 5. 机器人账号检测及其他相关任务 本数据集将按月更新。自2019年4月起,每份样本均遵循如下命名格式:german-tweet-sample-<YEAR>-<MONTH>.zip,单份样本大小约1GB。每份样本包含若干经gzip压缩的JSON格式记录文件,每批记录约包含15万条推文/账号数据,单批大小约18MB。



