adaption-mental-health-counseling
收藏资源简介:
该数据集是一个优化的心理健康咨询对话数据集,基于原始数据通过Adaption AutoScientist平台进行重构和增强。它包含34,529个高质量的一对一对话样本,采用提示-完成配对结构,记录了寻求帮助的个人与持证心理健康专业人士之间的交流。对话内容涵盖创伤、关系问题、焦虑、抑郁等敏感话题,并提供了富有同理心、情境感知的专业回应。数据集主要用于指令微调,旨在训练语言模型以更好地处理医疗和心理咨询场景中的支持性对话。质量评级为B级,相对原始数据质量提升78.0%。领域分布以医疗(66%)为主,辅以约会(24%)和个人成长(6%)。所有对话均为英文,语气风格以有帮助的(46%)、共情的(38%)和深思熟虑的(8%)为主。该数据集是Adaption AutoScientist Challenge 2026项目的一部分,通过自动化的质量评估、数据清洗、指令示例优化和对话一致性增强等处理流程生成。
This dataset is an optimized mental health counseling dialogue dataset, reconstructed and enhanced from the original data via the Adaption AutoScientist platform. It contains 34,529 high-quality one-on-one dialogue samples, structured in prompt-completion pairs, documenting exchanges between individuals seeking help and licensed mental health professionals. The dialogues cover sensitive topics such as trauma, relationship issues, anxiety, and depression, providing empathetic, context-aware professional responses. The dataset is primarily used for instruction fine-tuning, aiming to train language models to better handle supportive dialogues in medical and psychological counseling scenarios. It has a quality rating of B, with a 78.0% improvement over the original data quality. The domain distribution is dominated by medical (66%), supplemented by dating (24%) and personal growth (6%). All dialogues are in English, with tone styles primarily helpful (46%), empathetic (38%), and thoughtful (8%). This dataset is part of the Adaption AutoScientist Challenge 2026 project, generated through automated processing workflows including quality assessment, data cleaning, instruction example optimization, and dialogue consistency enhancement.
数据集概述
- 数据集名称: adaption-mental_health_counseling
- 语言: 英语(100%)
- 多语言性: 单语
- 数据集规模: 10K < n < 100K(具体为 34,529 条数据点)
- 来源数据集: Amod/mental_health_counseling_conversations
- 数据集类型: 指令微调数据集,由提示-回答对构成
内容与领域
该数据集包含寻求帮助的个人与持证心理健康专业人员之间的一对一高质量对话,涵盖创伤、人际关系问题、焦虑和抑郁等敏感话题,回应具有共情性和上下文感知能力。适用于微调语言模型以处理医疗和咨询场景中的支持性对话。
领域分布:
- 医疗(66%)
- 约会(24%)
- 个人成长(6%)
语气分布:
- 有帮助(46%)
- 共情(38%)
- 深思熟虑(8%)
数据质量与处理
- 最终质量等级: B
- 相对质量提升: 78.0%
- 处理方法: 使用 Adaption AutoScientist 流水线进行重制,包括评估数据质量、清洗和优化数据、改进指令遵循示例、增强对话一致性,生成适合语言模型微调的适配数据集。
标签
- adaption
- instruction-tuning
- medical
- dating
- personal-growth
数据集规模
- 数据点数量: 34,529 条




