A reusable longitudinal EHR-derived dataset for breast cancer recurrence research
收藏资源简介:
This record contains the public-use release package associated with the manuscript “A reusable longitudinal EHR-derived dataset for breast cancer recurrence modeling and clinical informatics research”. The resource provides a longitudinal, real-world, electronic health record (EHR)-derived breast cancer dataset from the University Clinical Centre Maribor, Slovenia. The dataset is intended to support patient-level secondary analyses of breast cancer recurrence, real-world oncology data quality assessment, cohort characterization, clinical informatics research, and future recurrence prediction studies. The public-use dataset is provided as data.csv and contains one row per patient. It includes 2,064 patients and 39 columns. The main grouping variable is cohort, which classifies 115 patients as recurrent and 1,949 patients as non-recurrent. Recurrence status should be interpreted as observed recurrence within the available institutional EHR follow-up window. Patients labelled as non-recurrent had no observed recurrence in the available records; this does not imply confirmed lifetime absence of recurrence. The release also includes cleaned Jupyter notebooks documenting the data construction and analysis workflow, together with output tables describing the file inventory, dataset summary and outcome distribution, public variable inventory, missingness, recurrence summaries by surgery year, baseline characteristics by recurrence cohort, and re-identification risk screening. The public-use dataset is a reduced and generalized version of the internal analysis dataset. To reduce disclosure risk, direct identifiers and exact clinical dates were removed or generalized before release. A re-identification risk screening was performed before public release, including checks for direct identifiers, exact date variables, small categorical levels, selected quasi-identifier combinations, and cells below a practical threshold of k = 5. These outputs are included for transparency and should be interpreted as disclosure-risk documentation, not as a guarantee of zero re-identification risk. Users must not attempt to re-identify individuals. The dataset is not intended as a note-level natural language processing benchmark, because the released file contains aggregated patient-level variables rather than original clinical notes or note-level extraction outputs. The GitLab development repository is available at: https://gitlab.com/humadex/breast-cancer-recurrence-ehr-dataset The public dataset and generated documentation in the Zenodo archive are released under Creative Commons Attribution 4.0 International (CC BY 4.0). The analysis code and notebooks in the GitLab repository are released under the Apache License 2.0.



