LLM-Based Synthetic Dataset for Mental Health Text Analysis
收藏Mendeley Data2026-07-04 收录
官方服务:
资源简介:
This dataset contains a paired parallel corpus of authentic and synthetic mental health disclosures. The human corpus consists of 5,184 real Reddit posts collected from three mental health communities: Reddit r/depression, r/anxiety, and r/mentalhealth. Each post was rewritten under standardized conditions by multiple contemporary large language models, producing a synthetic corpus designed for comparative analysis of linguistic, emotional, and stylistic variation between human and AI-generated mental health discourse. The dataset supports research in computational linguistics, machine behavior, AI psychology, and authenticity in computer-mediated mental health communication.
创建时间:
2026-07-03



