遇见数据集

Synthetic Chronic Kidney Disease (CKD) Audit Dataset - 200 Encounters for QA Training and EHR Coding Validation

收藏
Zenodo2025-08-13 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains a synthetic, fully de-identified electronic health record (EHR) sample representing 200 encounters for patients with chronic kidney disease (CKD). It was generated for quality assurance (QA) training, coding validation exercises, and workflow audits in healthcare data environments. The dataset mimics realistic variations in documentation, coding patterns, and lab values found in real-world EHR systems, but contains no real patient information. Data elements include: Demographics (Patient ID, Encounter ID, Coder name – all fictional) Clinical coding (ICD-10 CKD stage codes, diabetes & hypertension codes, combination codes) Laboratory values (eGFR, ACR) with dates Comorbidity indicators (diabetes, hypertension, dialysis, transplant status) Free-text note snippets with stage descriptions and linkage phrases The dataset was intentionally seeded with plausible inconsistencies to support QA scenarios: Under-specified CKD codes (e.g., N18.30) Missing combination codes for diabetes/HTN with CKD Dialysis status without N18.6 Lab-to-code mismatches (e.g., ESRD with high eGFR) Duplicate or missing secondary diagnoses Intended uses: Practice running automated data audits Train AI or rule-based systems to flag documentation/coding discrepancies Demonstrate data cleaning, merging, and report generation workflows Teach ICD-10 CKD coding guidelines in simulated settings Important:This is a synthetic dataset created entirely by AI-assisted generation and manual scenario design. It does not contain any identifiable or actual patient information.

本数据集包含一份完全去标识化的合成电子健康记录(Electronic Health Record, EHR)样本,涵盖慢性肾脏病(Chronic Kidney Disease, CKD)患者的200次就诊记录。该数据集专为医疗数据环境下的质量保证(Quality Assurance, QA)培训、编码验证练习及工作流审计工作而生成。 本数据集模拟了真实电子健康记录系统中常见的文档书写、编码范式与实验室检测值的真实变异特征,但未包含任何真实患者信息。数据元素包括: - 人口统计学信息(患者ID、就诊ID、编码员姓名——均为虚构内容) - 临床编码(国际疾病分类第十版(International Classification of Diseases 10th Revision, ICD-10)慢性肾脏病分期编码、糖尿病与高血压编码、组合编码) - 带日期的实验室检测指标(估算肾小球滤过率(estimated glomerular filtration rate, eGFR)、尿白蛋白/肌酐比值(albumin-to-creatinine ratio, ACR)) - 合并症指标(糖尿病、高血压、透析状态、移植状态) - 带分期描述与关联短语的自由文本笔记片段 本数据集有意植入了符合临床逻辑的不一致场景,以支撑质量保证相关演练: - 表述不明确的慢性肾脏病编码(如N18.30) - 缺失糖尿病/高血压合并慢性肾脏病的组合编码 - 仅标注透析状态但未使用N18.6编码(终末期肾脏病编码) - 实验室指标与编码不匹配(如标注终末期肾脏病(End-Stage Renal Disease, ESRD)但估算肾小球滤过率处于较高水平) - 重复或缺失的次要诊断 预期用途: - 演练自动化数据审计流程 - 训练人工智能或基于规则的系统,以识别文档记录与编码的不符之处 - 演示数据清洗、合并与报告生成的工作流 - 在模拟环境中教授国际疾病分类第十版慢性肾脏病编码规范 重要提示:本数据集完全通过人工智能辅助生成与人工场景设计构建,未包含任何可识别的真实患者信息。

提供机构:
Zenodo
创建时间:
2025-08-13
二维码
社区交流群
二维码
科研交流群
商业服务