Synthetic whole-exome sequencing data of virtual cancer patients (Clinical tables and metadata)
收藏资源简介:
Sharing genetic data from cancer patients is vital for developing improved diagnostic tools and training artificial intelligence models in medicine. However, given that genetic information is inherently unique to each individual, public data sharing entails substantial privacy risks. To address this challenge, this project has constructed a dataset of 'virtual cancer patients.' Specifically, we generated artificial genomic profiles that closely mimic real cancer data in both appearance and behavior, yet contain no actual patient-derived information. These profiles were produced by learning from a limited set of real clinical reports, validated through medical rule-based checks and artificial intelligence to ensure biological plausibility, and subsequently used to simulate raw sequencing reads by inserting the virtual mutations into a widely adopted public reference genome. This dataset enables researchers worldwide to freely test cancer analysis pipelines and train machine learning models without compromising real patient genomic privacy, thereby accelerating the advancement of precision medicine.



