Dataset and Analyses for Extracting Schemas from Thought Records using Natural Language Processing
收藏资源简介:
This dataset contains all data and analysis scripts pertaining to the research conducted for the PLOSOne paper: “Natural language processing for cognitive therapy: extracting schemas from thought records.” The cognitive approach to psychotherapy aims to change patients' maladaptive schemas, that is, overly negative views on themselves, the world, or the future. To obtain awareness of these views, they record their thought processes in situations that caused pathogenic emotional responses. To date, the schemas underlying such thought records have been largely manually identified. Using recent advances in natural language processing, we take this one step further by automatically extracting schemas from thought records. We used the Amazon Mechanical Turk crowd sourcing platform to collect a set of 1600 thought records. In total, these thought records contain 5747 thoughts of various depth levels, with the automatic thought constituting the most shallow level and the core belief the deepest level. We here deliver: 1. a natural language dataset: the thoughts delineated by participants in the scenario-based and open thought records2. reliability analyses: all thoughts were labeled with respect to the degree to which they reflect a set of 9 possible schemas by the first author. An independent second coder also labeled a sample of the thoughts.3. analyses to determine whether automatic identification of thoughts is possible.4. additional materials (scenarios, instruction videos, qualtrics survey, osf preregistration form) that could assist in the replication of the study.
本数据集涵盖为发表于《公共科学图书馆·综合》(PLOS ONE)的论文《用于认知疗法的自然语言处理(Natural Language Processing, NLP):从思维记录中提取认知图式(schemas)》开展的相关研究的全部数据与分析脚本。认知疗法作为心理治疗的重要范式,旨在改变患者的适应不良认知图式——即对自身、周遭世界或未来所持有的过度消极认知。为使患者意识到此类认知偏差,治疗师会指导患者在引发病理性情绪反应的情境下记录自身的思维过程。迄今为止,此类思维记录背后的认知图式主要依赖人工识别。借助近期自然语言处理领域的研究进展,本研究实现了从思维记录中自动提取认知图式的突破。本研究借助亚马逊众包平台(Amazon Mechanical Turk)收集了1600条思维记录,总计涵盖5747条不同深度层次的思维内容,其中自动思维为最浅层内容,核心信念为最深层内容。本数据集提供如下内容:1. 自然语言数据集:参与者在情境式与开放式思维记录中标注界定的思维内容;2. 信度分析:第一作者针对所有思维内容是否符合9种预设认知图式的程度进行了标注,同时另有一名独立编码员对部分思维样本开展了标注工作;3. 可行性分析:用于验证思维内容自动识别可行性的相关分析;4. 辅助研究复现的附加材料:包括情境脚本、指导视频、Qualtrics调查问卷以及开放科学框架(Open Science Framework, OSF)预注册表单。




