CODIPAS: Cognitive Distortions In Patient Speech
收藏资源简介:
CODIPAS is a dataset designed to support the detection and recognition of cognitive distortions in patient speech. Each instance in the dataset represents a fragment extracted from a patient’s question, annotated to indicate whether it contains any cognitive distortions and, if so, which specific types are present. It builds upon the Cognitive Distortion Detection dataset (Kaggle) by Shreevastava & Foltz, which originally selected and annotated patient questions from the Therapist Q&A dataset, a collection of questions and responses from an online forum where individuals sought psychological advice from certified therapists. This resource provides potential examples of emotionally charged and cognitively biased language. In CODIPAS, however, the entire process was redesigned and re-executed to achieve higher annotation quality and finer granularity. Expert psychologists re-examined all patient questions and re-identified distorted parts from scratch, allowing each question to be segmented into multiple fine-grained fragments when relevant. Every fragment was then annotated through an expert consensus process, ensuring consistent and reliable labeling of cognitive distortions. In addition, new instances were generated using generative AI to improve class balance and diversity. These synthetic examples were first validated by clinical experts for inclusion and then annotated by psychologists following the same protocol as the original data. Data Structure Each instance in CODIPAS includes: a new unique identifier associated with each Distorted Part the full text of the patient statement the extracted distorted fragment expert-assigned labels for Dominant Distortion and Secondary Distortion (or No distortion) This structure allows CODIPAS to be used for both binary (distorted vs. non-distorted) and multiclass classification tasks, providing a robust benchmark for studying the linguistic patterns of cognitive biases. Types of Cognitive Distortion Labels included: All-or-nothing thinking Overgeneralization Mental filter Should statements Labeling Personalization Magnification Emotional reasoning Mind reading Fortune-telling No distortion: indicates that the fragment does not contain any identifiable cognitive distortion.



