遇见数据集

Synthetic Generation from Real LASIK Procedures (TENEO 317 model 1 / ZYWAVE 3 / ORBSCAN 3, 2017–2024)

收藏
Zenodo2026-04-21 更新2026-05-29 收录
官方服务:

资源简介:

DATASET TITLE TabDDPM-Generated Synthetic LASIK Biometric Dataset for Refractive Surgery AI Model Development SUMMARY This dataset comprises 25,000 synthetic patient records generated via Tabular Denoising Diffusion Probabilistic Model (TabDDPM) from a real clinical cohort of 2,448 LASIK procedures performed between July 2017 and August 2024. The synthetic data preserves the joint distribution and multivariate dependencies of real biometric measurements while ensuring strict patient anonymity and topological novelty (volumetric expansion ratio = 1.03). CLINICAL CONTEXT Preoperative and postoperative biometric measurements were acquired using three complementary diagnostic modalities: - Corneal morphometry and scanning-slit elevation topography (ORBSCAN 3, n=67 features) - Ocular aberrometry via Shack-Hartmann wavefront sensing (ZYWAVE 3, n=40 features, cycloplegia) - Excimer laser console parameters (TENEO 317 Model 1/2, n=40 features) GENERATIVE PROCESS The synthetic cohort was generated via Bayesian-optimized TabDDPM to maximize statistical fidelity (KSD=0.05, AUC-ROC=0.55) while preventing memorization of source records (DCR mean=1.821, NNDR mean=0.908, 0% exact match rate). The model successfully synthesized 1,334 novel rare clinical cases beyond the empirical convex hull of the training distribution, validating volumetric expansion into biologically plausible but unobserved phenotypes. VALIDATION METRICS - Kolmogorov-Smirnov Distance (marginal fidelity): 0.05 - Adversarial discriminator AUC-ROC (structural realism): 0.55 - Train-on-Synthetic, Test-on-Real (TSTR) mean R²: 0.91 (vs. baseline TRTR R²=0.90) - Distance to Closest Record (privacy): mean DCR=1.821 - Nearest Neighbor Distance Ratio (topological separation): mean NNDR=0.908 PRIVACY & COMPLIANCE This dataset is compliant with GDPR and French data protection regulations (Jardé Law, CNIL Reference Methodology MR-004). No real patient identifiers are present. Synthetic data generation and topological filtering ensure that no source records can be reconstructed or re-identified. DATA STRUCTURE Each record (n=2448; n=25,000; n=95000) contains: - 12 input features: pre-op/post-op aberrometric states, surgical zone parameters - 4 target labels: planned console inputs (Q-factor, defocus, astigmatism J0/J45) - All values denormalized to original clinical scale (Diopters, mm, µm RMS, dimensionless) DATA TRANSFORMATION & PREPROCESSING All synthetic records underwent the following standardized transformations prior to release: - Q-Factor (QVAR) values: uniform offset of +0.66 units added to all instances. This offset is documented to ensure reproducibility when training external models on this dataset. INTENDED USE This dataset is provided as an open-access benchmark for: 1. Model training in data-scarce refractive surgery AI applications 2. Validation of synthetic data generation quality in medical imaging 3. Development of transfer learning architectures for ophthalmic prediction tasks 4. Reproducible research in AI-assisted surgical planning RELATED PUBLICATIONS Garnier S. Clinical Safety and Data Sovereignty of AI in Laser Refractive Surgery: A Foundation Model Framework with Synthetic Benchmarking, Transfer Learning and Conformal Prediction. Journal of Refractive Surgery. [Submitted 2026]. COMPUTATIONAL REPRODUCIBILITY Generation scripts and hyperparameters are available at GitHub [youllfindzelink]. TabDDPM implementation: Kotelnikov et al., TabDDPM: modelling tabular data with diffusion models. Proc Mach Learn Res. 2023;202:17564-17579. LICENSE Creative Commons Attribution 4.0 International (CC-BY-4.0)

提供机构:
Zenodo
创建时间:
2026-04-21
二维码
社区交流群
二维码
科研交流群
商业服务