遇见数据集

Monthly Samples of German Tweets

收藏
Zenodo2023-01-12 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This dataset contains German tweets and Twitter accounts recorded from the public Twitter Streaming API (via Tweepy) using the following filters: terms: <em>'a'</em>, <em>'e'</em>, <em>'i'</em>, <em>'o'</em>, <em>'u'</em>, and <em>'n'</em> language: <em>'de'</em> This filter combination should record (almost) all German tweets (in German it is very unlikely that terms do not contain vowels or the frequently used character <em>'n'</em>). This dataset might be useful for the following use cases: Natural language processing (focussing on Twitter specifics in German, there exist only little datasets focussing this specific aspect) Social Network Analysis (Twitter network) Identifying behavioral patterns (retweeting, quoting, replying, hate speech, ...) Sharing political (or other domain-specific) content Bot detection and more ... This dataset will be updated monthly. Each sample (starting in April 2019) will follow the following naming pattern: german-tweet-sample-<em>&lt;YEAR&gt;</em>-<em>&lt;MONTH&gt;</em>.zip (size: ~ 1GB) It will contain several bunches of recorded JSON gzipped files. Each bunch of records contains approximately 25k recorded tweets/accounts (size: ~ 3MB).

本数据集收录了通过公开Twitter流式API(Twitter Streaming API)并借助Tweepy库抓取的德语推文与Twitter账号,抓取时采用了如下过滤规则:关键词为'a'、'e'、'i'、'o'、'u'及'n',语言限定为德语(语言代码'de')。 该过滤组合可覆盖(几乎)全部德语推文——在德语语境中,不含元音或高频字符'n'的推文几乎不存在。 本数据集可应用于以下研究场景: 1. 自然语言处理(Natural Language Processing,聚焦德语Twitter平台的专属特性,目前针对该细分方向的公开数据集较为稀缺) 2. 社交网络分析(针对Twitter网络结构) 3. 行为模式识别(包括转发、引用、回复、仇恨言论识别等) 4. 政治(或其他垂直领域)内容传播分析 5. 机器人账号检测 以及更多潜在应用场景…… 本数据集将按月更新。自2019年4月起,每份样本均遵循如下命名规范: german-tweet-sample-<YEAR>-<MONTH>.zip(文件大小约1GB) 每份压缩包内含多组经gzip压缩的JSON格式记录文件,每组记录约包含25000条推文或账号数据,单组文件大小约3MB。

提供机构:
Zenodo
创建时间:
2019-05-13
二维码
社区交流群
二维码
科研交流群
商业服务