遇见数据集

BharBharBinks/MentalChat16K

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含两个部分:1. Synthetic Data 10K:由9,775个合成对话组成,模拟咨询师与客户之间的对话,涵盖33个心理健康主题,如关系、焦虑、抑郁、亲密关系和家庭冲突等。这些对话使用OpenAI GPT-3.5 Turbo模型和定制化的Airoboros自生成框架生成,通过提供清晰的指令来生成患者查询,并确保主题比例以真实模拟人类治疗师-客户互动的复杂性和多样性。2. Interview Data 6K:由6,338个问答对组成,源自378个访谈转录,这些转录来自正在进行的行为干预会话的音频记录,由行为健康教练和临终关怀护理人员参与,并由人类专家转录。使用本地Mistral-7B-Instruct-v0.2模型对转录进行转述和总结,将每页转录转换为护理人员与行为健康教练之间的单轮对话,并过滤掉问题和回答中少于40词的对话。整个数据集旨在为语言模型提供广泛的心理健康状况和治疗策略的暴露,以促进更现实和有效的心理咨询对话。

This dataset consists of two parts: 1. Synthetic Data 10K: 9,775 synthetic conversations between a counselor and a client, covering 33 mental health topics such as Relationships, Anxiety, Depression, Intimacy, and Family Conflict. The conversations were generated using the OpenAI GPT-3.5 Turbo model and a customized adaptation of the Airoboros self-generation framework, with clear instructions for generating patient queries to authentically mimic the complexity and diversity of human therapist-client interactions. 2. Interview Data 6K: 6,338 question-answer pairs from 378 interview transcripts collected from an ongoing clinical trial, based on audio recordings of behavioral intervention sessions between behavior health coaches and caregivers of individuals in palliative or hospice care, transcribed by human experts. The local Mistral-7B-Instruct-v0.2 model was used to paraphrase and summarize the transcripts into single rounds of conversation, filtered for those with at least 40 words in both question and answer. The dataset aims to equip language models with exposure to psychological conditions and therapeutic strategies for realistic counseling conversations.

提供机构:
BharBharBinks
二维码
社区交流群
二维码
科研交流群
商业服务