遇见数据集

CODIPAD: Cognitive Distortions in Patient Discourse

收藏
Zenodo2026-05-20 更新2026-05-26 收录
官方服务:

资源简介:

CODIPAD is a dataset designed to support the detection and recognition of cognitive distortions in patient discourse. Each instance in the dataset represents a fragment extracted from a patient’s statement, annotated to indicate whether it contains any cognitive distortion and, if so, which specific type or types are present. It builds upon the Cognitive Distortion Detection dataset (Kaggle) by Shreevastava & Foltz, which originally selected and annotated patient questions from the Therapist Q&A dataset, a collection of questions and responses from an online forum where individuals sought psychological advice from certified therapists. This resource provides potential examples of emotionally charged and cognitively biased language. In CODIPAD, however, the entire process was redesigned and re-executed to achieve higher annotation quality and finer granularity. Expert psychologists re-examined all patient questions and re-identified distorted parts from scratch, allowing each question to be segmented into multiple fine-grained fragments when relevant. Every fragment was then annotated through an expert consensus process, ensuring consistent and reliable labeling of cognitive distortions. In addition, new instances were generated using generative AI to improve class balance and diversity. These synthetic examples were first validated by clinical experts for inclusion and then annotated by psychologists following the same protocol as the original data. Data Structure Each instance in CODIPAD includes: id_distorted_part: a unique identifier associated with each analyzed fragment. id_patient_question: an identifier for the full patient statement; this identifier is shared across instances derived from the same statement. origin: indicates whether the instance is organic or synthetic. patient_question: the full text of the patient statement. distorted_part: the extracted fragment analyzed for cognitive distortion. distortions: expert-assigned labels for Dominant Distortion and Secondary Distortion. The distortions object contains: dominant_distortion: the main cognitive distortion label assigned to the analyzed fragment, or No distortion when the fragment does not contain any identifiable cognitive distortion. secondary_distortion: an optional secondary cognitive distortion label, or null when no secondary distortion is assigned. This structure allows CODIPAD to be used for both binary (distorted vs. non-distorted) and multiclass classification tasks involving specific cognitive distortion categories, providing a robust benchmark for studying the linguistic patterns of cognitive biases. The "origin" field also enables analyses that distinguish between organic and synthetic instances. Types of Cognitive Distortion Labels included: All-or-nothing thinking Overgeneralization Mental filter Should statements Labeling Personalization Magnification Emotional reasoning Mind reading Fortune-telling No distortion: indicates that the fragment does not contain any identifiable cognitive distortion.

提供机构:
Zenodo
创建时间:
2026-05-20
二维码
社区交流群
二维码
科研交流群
商业服务