遇见数据集

LLM-Based Synthetic Dataset for Mental Health Text Analysis

收藏
Mendeley Data2026-07-04 收录
官方服务:

资源简介:

This dataset contains a paired parallel corpus of authentic and synthetic mental health disclosures. The human corpus consists of 5,184 real Reddit posts collected from three mental health communities: Reddit r/depression, r/anxiety, and r/mentalhealth. Each post was rewritten under standardized conditions by multiple contemporary large language models, producing a synthetic corpus designed for comparative analysis of linguistic, emotional, and stylistic variation between human and AI-generated mental health discourse. The dataset supports research in computational linguistics, machine behavior, AI psychology, and authenticity in computer-mediated mental health communication.

创建时间:
2026-07-03
二维码
社区交流群
二维码
科研交流群
商业服务