A large-scale Twitter dataset for drug safety applications mined from publicly existing resources
收藏资源简介:
This dataset consists of 1,181,993 Tweet Ids, obtained as a result from the paper - A large-scale Twitter dataset for drug safety applications mined from publicly existing resources. The Tweet Ids can be hydrated using <strong>twarc</strong>, a command line tool and Python library for archiving Twitter JSON data. Please follow these instructions to install twarc - https://github.com/DocNow/twarc It will take less than 4 hours to hydrate the tweet ids using twarc. If you intend to use the dataset, please cite the paper.
本数据集共包含1,181,993条推文标识符(Tweet ID),其源自论文《从公开资源中挖掘得到的面向药物安全应用的大规模Twitter数据集》。该数据集的推文标识符可通过twarc完成水化处理:twarc是一款用于归档Twitter JSON数据的命令行工具及Python库。请按照以下说明安装twarc:https://github.com/DocNow/twarc。使用twarc完成推文水化操作的耗时不足4小时。若您计划使用本数据集,请引用该论文。



