遇见数据集

CoCAST: Synthetic psychiatric questionnaire datasets for depression and anxiety

收藏
Zenodo2026-09-28 更新2026-10-01 收录
官方服务:

资源简介:

This release contains synthetic psychiatric questionnaire datasets accompanying the CoCAST study. CoCAST, short for Copula-Conditioning from Aggregate summaries for Synthesis of Tabular data, combines a statistical model of cross-disorder severity with a pretrained language model to generate questionnaire responses from aggregate cohort summaries and DSM-5 context. Each dataset contains responses to 69 questionnaire items covering depression, separation anxiety, specific phobia, social anxiety, panic, agoraphobia and generalized anxiety, together with synthetic age and recorded-sex variables. The collection supports research on synthetic tabular data, preservation of cross-domain symptom associations, downstream prediction and controlled data augmentation. The reference cohort was drawn from the first assessment wave of the study by Vidal-Arenas et al. (2025). The original participant-level dataset and associated study materials are available from the source study’s OSF repository. CoCAST uses aggregate summaries computed from this cohort, while CTGAN and TVAE were trained on its participant-level records. The original participant-level data are not included in this synthetic dataset release. The release includes CoCAST datasets generated with Qwen3.5-4B, Qwen3.5-9B and Qwen3.5-27B under natural and balanced severity distributions, a comparison of questionnaire-specific and full-profile severity context, and a copula ablation that samples severity targets independently while retaining the same marginal score distributions. CTGAN and TVAE datasets are included as record-trained comparators. Experiments use three generation seeds. There are 36 datasets containing 564 records each and six expanded Qwen3.5-27B pools containing 2,260 records each. Each expanded pool includes its corresponding 564-record dataset as a prefix, so these files should not be combined as independent samples. The accompanying data dictionary defines the variables, response ranges and missing-value conventions, while the dataset index identifies the generation conditions and relationships between files. The CoCAST implementation, including generation and evaluation code, configuration files and usage instructions, is available at: https://github.com/Adamjakobsen/Copula-Conditioning-from-Aggregate-summaries-for-Synthesis-of-Tabular-data

提供机构:
Zenodo
创建时间:
2026-09-28
二维码
社区交流群
二维码
科研交流群
商业服务