Human-Verified Clinical Reasoning Dataset for Trustworthy Medical AI
收藏资源简介:
本数据集由上海人工智能实验室等研究机构构建,包含31,247个医疗问题-答案对,每个问题-答案对都伴有专家验证的推理链(CoT)解释。数据集涵盖多个临床领域,通过可扩展的人类-LLM混合流程进行筛选和整理。LLM生成的推理链由医疗专家根据结构化评分标准进行迭代审查、评分和优化,确保输出的高质量和临床相关性。该数据集公开可用,为医疗LLM的开发提供了关键资源,旨在促进医疗领域安全、可解释的AI发展。
This dataset was constructed by research institutions including the Shanghai AI Laboratory. It comprises 31,247 medical question-answer pairs, each accompanied by expert-validated Chain-of-Thought (CoT) explanations. Spanning multiple clinical domains, the dataset was screened and curated via a scalable human-LLM hybrid workflow. The LLM-generated CoT explanations are iteratively reviewed, scored, and optimized by medical experts in accordance with structured scoring criteria, ensuring high output quality and clinical relevance. This publicly available dataset serves as a critical resource for the development of medical LLMs, with the aim of advancing safe and interpretable AI development within the healthcare sector.




