HealthDataHub/PARHAF
收藏资源简介:
PARHAF 是一个开放的法语语料库,包含由人类撰写的虚构患者临床报告。该数据集旨在在严格的健康数据保护约束下,支持临床自然语言处理(NLP)系统的开发和评估。每个患者记录都附有结构化的临床信息(如诊断、手术、护理路径、出院数据)。这些记录由资深医学住院医师撰写,并由同一专业的另一位资深医学住院医师审核。数据集包含4259名患者、6190份文档,总计3,952,583个单词,覆盖心脏病学、心血管外科、重症监护、消化外科、胃肠肝病学、普通内科、老年病学、妇科学、血液学、传染病学、内科学、肾病学、神经学、产科学、肿瘤学、骨科与创伤外科、病理学、儿科学、肺病学和泌尿学等多个医学专业。数据来源于法国国家医院索赔数据库(SNDS),通过采样观察分布构建临床场景,以减少常见病症的过度代表并包含较少见情况。预期用途包括共享临床笔记和注释、汇集临床NLP社区努力、基准测试法语医学LLM、支持可重复的临床NLP研究、促进医学教学、推动PARTAGES项目的7个用例,以及实现隐私安全的数据增强。排除用途包括临床决策或患者护理、临床验证或性能声明、推广到未见过的医院或实践、流行病学推断、评估真实世界安全风险、替代真实临床数据进行部署,以及压力测试模型在真实临床语言上的表现。数据集仅限医院临床文档,聚焦核心文档类型,不包括影像报告、处方、转诊信等类别。
PARHAF is an open French corpus of human-authored clinical reports of fictional patients. It was created to support the development and evaluation of clinical NLP systems under strict health-data protection constraints. Each patient is each documented with structured clinical information (diagnosis, procedures, care pathway, discharge data). Each patient record was written by a senior medical resident and reviewed by another senior medical resident, from the same specialty. The dataset includes 4259 patients, 6190 documents, and 3,952,583 words across specialties such as Cardiology, Cardiovascular Surgery, Critical Care, Digestive Surgery, Gastro-Hepatology, General Internal Medicine, Geriatrics, Gynecology, Hematology, Infectious Diseases, Internal Medicine, Nephrology, Neurology, Obstetrics, Oncology, Orthopedic & Trauma Surgery, Pathology, Pediatrics, Pulmonology, and Urology. Data origin is from the French national hospital claims database (SNDS), with scenarios built by sampling observed distributions to reduce over-representation of common conditions and include less frequent situations. Intended usages include sharing clinical notes and annotations, pooling efforts within the clinical NLP community, benchmarking French medical LLMs, enabling reproducible clinical NLP research, supporting medical teaching, promoting PARTAGES 7 use cases, and enabling privacy-safe data augmentation. Excluded usages cover clinical decision-making or patient care, clinical validation or performance claims, generalization to unseen hospitals or practices, epidemiological inference, assessing real-world safety, replacing real clinical data for deployment, and stress-testing models on realistic clinical language. The dataset is limited to hospital-based clinical documentation and focuses on selected document types, excluding categories like imaging reports and prescriptions.




