VRCardio-Helper-Database: natural-language query benchmark and dataset documentation
收藏资源简介:
Reproducibility materials accompanying the paper "A Multimodal ECG-Centric Clinical Data Platform with AI-Assisted Querying for Cardiovascular Research." This record contains the natural-language query (NLQ) benchmark (40 clinically motivated queries with consensus ground-truth SQL, robustness reformulation clusters, and per-query precision/recall; overall precision 0.90, recall 0.88, robustness IoU 0.97), the schema/data dictionary for the integrated repository's derived tables (~1,185,806 ECG records from ~432,338 patients across nine sources), a per-source provenance and licensing manifest, and de-identified aggregate (population-level) summary statistics. It does not redistribute the per-record clinical data or the raw ECG signals: MIMIC-IV-ECG is governed by the PhysioNet Credentialed Health Data License; the other public datasets must be obtained from their original providers; and the in-house VR-CARDIO cohort (46 recordings) is governed by its clinical-study ethics approval and informed consent and is available from the corresponding author on reasonable request. The integration code is published in the companion GitHub repository (see Related works). Interval metrics (PR, QRS, QTc) are affected by a delineation artifact described in the data dictionary and should be cross-checked against the SCP-ECG diagnostic codes.



