深圳市宝安区人民医院阿尔兹海默症临床文本数据集
收藏资源简介:
本数据集为阿尔茨海默症住院患者真实世界临床诊疗文本数据集,聚焦于疑似或确诊阿尔茨海默病患者的完整住院诊疗过程。本数据集涵盖了2021年12月31日至2025年12月1日期间,深圳市宝安区人民医院神经内科、老年科及综合内科等科室收治的约260例患者的原始临床记录,包括入院记录、病程记录、手术信息、出院小结等非结构化/半结构化中文文本,以及对应的脱敏患者ID、入院日期、检查检验结果等结构化字段。在此基础上,生成了以患者为中心的多元化临床数据。通过分析评估阿尔茨海默病患者在真实住院场景中的症状描述演化模式、共病记录习惯及诊疗叙事结构等维度,可帮助深入理解中文医疗语境下神经认知障碍的临床表征规律,为构建面向中文临床文本的自然语言处理基准、疾病表型挖掘方法及真实世界疾病描述体系提供高质量语料支撑,进而推动人工智能在老年神经退行性疾病观察性研究与知识发现中的应用。
This dataset is a real-world clinical diagnosis and treatment text dataset for hospitalized patients with Alzheimer's disease, focusing on the complete inpatient diagnosis and treatment course of patients with suspected or confirmed Alzheimer's disease. It covers original clinical records of approximately 260 patients admitted to departments including Neurology, Geriatrics, and General Internal Medicine of Shenzhen Bao'an District People's Hospital from December 31, 2021 to December 1, 2025. These records include unstructured and semi-structured Chinese clinical texts such as admission notes, progress notes, surgical information, and discharge summaries, as well as corresponding structured fields including de-identified patient IDs, admission dates, laboratory and examination results, etc. On this basis, patient-centric diversified clinical data were generated. By analyzing and evaluating dimensions such as the evolutionary patterns of symptom descriptions, comorbidity recording habits, and diagnostic and treatment narrative structures of Alzheimer's disease patients in real inpatient scenarios, this dataset can help gain in-depth insights into the clinical manifestation rules of neurocognitive disorders in the context of Chinese medical practice. It provides high-quality corpus support for building natural language processing benchmarks for Chinese clinical texts, disease phenotype mining methods, and real-world disease description systems, thereby promoting the application of artificial intelligence in observational studies and knowledge discovery of age-related neurodegenerative diseases.




