De-identified Reddit Textual Corpora for Depression and Male Health (LOH) NLP Analysis
收藏资源简介:
This repository provides a specialized collection of de-identified textual data extracted from Reddit, curated for Natural Language Processing (NLP) research. The archive includes two distinct datasets focused on mental health and endocrinology. 1. Depression Corpus: A collection of discourse from r/depression, filtered through a multi-domain keyword strategy aligned with clinical scales (such as PHQ-9 and BDI-II).2. LOH (Loss of Health) Corpus: A dataset focused on male-specific health transitions and androgen deficiency, utilizing search parameters based on the AMS (Aging Males' Symptoms) lexicon. To protect user privacy, all identifiers were removed at the point of acquisition. The data was captured on December 31, 2025. These files (LOHNLP.docx and DepressionNLP.docx) serve as the source material for semantic co-occurrence network analysis. For the complete theoretical framework, methodology, and research findings, please refer to the associated peer-reviewed publication.



