遇见数据集

Monthly Samples of German Tweets

收藏
Zenodo2020-07-30 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This dataset contains German tweets and Twitter accounts recorded from the public Twitter Streaming API using the following filters: terms: <em>'a'</em>, <em>'e'</em>, <em>'i'</em>, <em>'o'</em>, <em>'u'</em>, and <em>'n'</em> language: <em>'de'</em> This filter combination should record (almost) all German tweets (in German it is very unlikely that terms do not contain vowels or the frequently used character <em>'n'</em>). This dataset might be useful for the following use cases: Natural language processing (focussing on Twitter specifics in German, there exist only little German datasets) Social Network Analysis (Twitter network) Identifying behavioural patterns (retweeting, quoting, replying, hate speech, ...) Sharing political (or other domain-specific) content Bot detection and more ... This dataset will be updated monthly. Each sample (starting in April 2019) will follow the following naming pattern: german-tweet-sample-<em>&lt;YEAR&gt;</em>-<em>&lt;MONTH&gt;</em>.zip (size: ~ 1GB) It will contain several bunches of recorded JSON gzipped files. Each bunch of records contains approximately 50k recorded tweets/accounts (size: ~ 6MB).

本数据集包含通过公开Twitter流式API (Twitter Streaming API)采集的德语推文与Twitter账号,采集时采用了如下过滤规则:检索词为字母‘a’、‘e’、‘i’、‘o’、‘u’及‘n’,语言过滤条件设为‘de’(德语)。该过滤组合可采集(几乎)全部德语推文——在德语语境中,不含元音或高频字符‘n’的词汇极为罕见。 本数据集可应用于以下场景:聚焦德语推文平台特性的自然语言处理(当前德语领域相关数据集仍较为稀缺)、社交网络分析(Twitter社交网络)、行为模式识别(涵盖转发、引用、回复、仇恨言论等)、政治(或其他特定领域)内容传播、机器人账号检测及其他相关领域。 本数据集将按月更新。自2019年4月起,每份样本均遵循如下命名格式:german-tweet-sample-<YEAR>-<MONTH>.zip(单份大小约1GB),其内包含多组经gzip压缩的JSON文件。每组记录约包含5万条推文/账号数据,单组大小约6MB。

提供机构:
Zenodo
创建时间:
2019-11-01
二维码
社区交流群
二维码
科研交流群
商业服务