Pseudonymized Complete Blood Count Test Data
收藏资源简介:
This dataset contains pseudonymized routine laboratory data collected at Klinikum Lippe Detmold, a German tertiary care center. The original dataset was collected over a period of 1,585 days and included a broad range of laboratory measurements. The version published here consists exclusively of complete blood count (CBC) analytes from 523,844 samples across 77,355 patients. It has been pre-filtered to exclude samples from patients with documented blood transfusions in the preceding 90 days, patients with only a single recorded sample, and samples lacking complete CBC analyte measurements. Each record in the dataset includes the following columns: SampleNum: pseudonymized sample identifier PatientNum: pseudonymized patient identifier Timestamp: sample collection timestamp (date-shifted for privacy) WardNum: enumerated code representing the hospital ward ERY: red blood cells (×10⁶ /μL) HK: hematocrit (%) LEUKO: white blood cells (×10³ /μL) HB: hemoglobin (g/dL) PLT: platelets (×10³ /μL) MCV: mean corpuscular volume (fL) MCHC: mean corpuscular hemoglobin concentration (g/dL) MCH: mean corpuscular hemoglobin (pg) RDW: red blood cell distribution width (%) This dataset was used as the basis for the study “Analyte Importance Analysis in Machine Learning-Based Detection of Wrong Blood in Tube (WBIT) Errors” and is shared to support reproducibility and further research in machine learning applications in laboratory medicine.



