遇见数据集

Long-CoDe: A Longitudinal Twitter Dataset of Depression in the COVID Era

收藏
Zenodo2023-01-16 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

# Long-CoDe: A Longitudinal Twitter Dataset of Depression in the COVID Era ## Content of Long-CoDe Long-CoDe contains 15 million tweets from 597 depressed users and 804 control users. The temporal span of collected tweets is from Jan. 1st, 2019 to Dec. 31, 2022. We do not have the right to release text of the tweets due to privacy issues. We include user-ids, tweet-ids, date of posting, and label of the users in this dataset. We disguised user-ids with C00n and D00n to protect privacy and to represent their group labels. Tweets can be accessed with tweet-ids from the Twitter-APIs. The format of the data is like follows: | user-id | tweet-id | label |<br> | ------- | -------- | -------------- |<br> | C001| xxxxxxxx | 0 (Control) |<br> | D001| xxxxxxxx | 1 (Depression) | ## Quality All users in both depressed and control groups posted steadily from 2019 to 2022. Control users were selected from 43 trending topics in 2020 evenly from Feb. to Dec. with broad interests. ## Structure The Long-CoDe dataset is a single CSV file with tabular data. The format of the CSV file is as follows: | user_id | tweet_id | label<br> | ------- | -------- | -------------- |<br> | C001| xxxxxxxx | 0 (Control) |<br> | D001| xxxxxxxx | 1 (Depression) | ## Potential uses of the dataset General analysis of tweets from the labeled depressed users, compared with the control.<br> Benchmark binary user-classification with contents of tweets or other extracted features.<br> Analyze the impact of Covid-19 on both depression and control groups in multiple phases of the pandemic. ## Data Collection The data collection procedure contains four steps: (a) identifying depressed users, (b) selecting control users, (c) collecting tweets in 2020 and identifying regular users, and (d) collecting tweets from 2019 to 2022 for all selected users. ### Identifying Depressed Users To identify targeted depressed users, we crawled tweets containing depression self-claims, such as ”I am/was/have been diagnosed with depression”, using regular expressions from Feb 2020 to Dec 2020,when Covid-19 had a significant influence on our society. The data was manually annotated, and keep the users who did a valid self-claim on depression diagnosis. ### Identifying Control User In an attempt to create a diverse control user population, we gathered users that have shown interests in a wide array of trending topics, so that this would result in users with different backgrounds. From that, we randomly selected 6,200 unique users, and filtered out the users who ever made depression claims to give us a total of 5,929 control users. ### Identifying Regular Users For each user from both groups, we crawled the tweets (without re-tweets) from Feb. to Dec. in 2020 and analyzed their posting behaviors. We keep the users who have stable posting behaviors. (posting between 75 and 205 tweets per month) ### Collecting tweets from 2019 to 2022 We crawled all tweets for each of the identified regular users for four years from Jan., 2019 to Dec., 2022.<br>

# Long-CoDe:新冠疫情时代抑郁症纵向推特数据集 ## Long-CoDe数据集内容 Long-CoDe数据集包含来自597名抑郁症用户与804名对照用户的1500万条推文。所采集推文的时间跨度为2019年1月1日至2022年12月31日。出于隐私保护原因,我们无权发布推文的原始文本。本数据集仅包含用户ID、推文ID、发布日期以及用户标签。我们使用C00n与D00n对用户ID进行匿名化处理,以保护用户隐私并区分其组别标签。可通过推文ID调用推特应用程序编程接口(Twitter API)获取原始推文。数据格式如下: | user-id | tweet-id | label | | ------- | -------- | -------------- | | C001| xxxxxxxx | 0 (对照) | | D001| xxxxxxxx | 1 (抑郁症) | ## 数据集质量 抑郁症组与对照组的所有用户均在2019年至2022年间保持稳定的推文发布频率。对照组用户于2020年2月至12月期间,从43个涵盖广泛兴趣领域的热门话题中均匀选取。 ## 数据集结构 Long-CoDe数据集为单个逗号分隔值(CSV)格式的表格数据文件。CSV文件格式如下: | user_id | tweet_id | label | | ------- | -------- | -------------- | | C001| xxxxxxxx | 0 (对照) | | D001| xxxxxxxx | 1 (抑郁症) | ## 数据集潜在应用场景 1. 针对标注为抑郁症用户与对照用户的推文开展通用分析 2. 基于推文内容或其他提取特征开展二分类用户分类任务的基准测试 3. 分析新冠疫情不同阶段对抑郁症组与对照组用户群体的影响 ## 数据采集流程 数据采集流程包含四个步骤:(a) 抑郁症用户识别,(b) 对照用户选取,(c) 2020年推文采集与常规用户筛选,(d) 为所有选中用户采集2019年至2022年的推文。 ### 抑郁症用户识别 为识别目标抑郁症用户,我们于2020年2月至12月(新冠疫情对社会产生显著影响的时期),通过正则表达式爬取包含抑郁症自我声明的推文,例如"我被诊断出患有抑郁症(I am/was/have been diagnosed with depression)"。随后对数据进行人工标注,保留那些对抑郁症诊断做出有效自我声明的用户。 ### 对照用户识别 为构建多样化的对照用户群体,我们收集了对各类热门话题表现出兴趣的用户,以覆盖不同背景的人群。随后从中随机选取6200名独立用户,并过滤掉曾发布过抑郁症相关声明的用户,最终得到5929名对照用户。 ### 常规用户筛选 针对两组中的每位用户,我们爬取了其2020年2月至12月期间的推文(不含转发推文),并分析其发布行为。我们保留了发布行为稳定的用户,即月均发布推文75至205条的用户。 ### 2019年至2022年推文采集 我们为所有筛选出的常规用户爬取了2019年1月至2022年12月这四年间的全部推文。

提供机构:
Zenodo
创建时间:
2023-01-16
二维码
社区交流群
二维码
科研交流群
商业服务