遇见数据集

GAP: Genetic Architecture for Phenotype Prioritization Framework for Diagnostic Gain of Long-Read Sequencing

收藏
Zenodo2026-08-10 更新2026-08-13 收录
官方服务:

资源简介:

Short-read sequencing leaves most rare-disease patients undiagnosed. Long-read sequencing resolves some of those cases but costs more, so patient selection matters. GAP scores every gene for the probability that long reads produce a diagnosis short reads missed, conditional on that gene being causal and short reads having been negative. Three routes are scored separately and combined on the hazard scale, because they are independent causes of rescuability rather than three measurements of one quantity: damaging positions inside sequence short reads cannot map or call (128,120 positions across 1,430 genes); ClinVar pathogenic and likely-pathogenic structural variants of 50 bp to 10 kb (3,901 variants across 1,037 genes); and pathogenic tandem-repeat loci from STRchive, graded by evidence (77 genes). The routes overlap far less than chance predicts (61 genes observed against 148 expected, odds ratio 0.35), which is why they are summed rather than averaged. Route weights are rescue-mechanism counts taken from a systematic review and individual-patient-data meta-analysis of long-read sequencing in short-read-negative patients. Everything is computed for four short-read to long-read pairs (Illumina-100 or Illumina-250 against ONT or PacBio HiFi), so the per-gene result shows which sequencing change would resolve each gene. Two outputs are included. A known-phenotype view ranks 11,918 HPO terms and 431 PanelApp panels by the lower bound of a one-sided 95% confidence interval on the conditional odds ratio. A dark-gene view lists 271 established disease genes carrying damaging positions inside short-read-dark sequence, 215 of them on live clinical PanelApp panels; for 211 of these, ClinVar holds no structural variant at all, so the dark-region route is the only signal. The deposit also contains an external check of the structural-variant route against expert curation. Deletion/duplication shares were extracted from 834 of 867 "Molecular Genetic Testing" tables across 1,011 GeneReviews chapters. ClinVar reproduces the expert ordering (Spearman rho = 0.43, 95% CI 0.36 to 0.50) but compresses the magnitude roughly eight- to tenfold, and the shortfall grows with how copy-number-driven the disease is. Reference set: GRCh38, GENCODE v49, MANE v1.5, ClinVar 20 May 2026, HPO 23 June 2026, STRchive 4 August 2026. The pipeline reproduces every file byte-for-byte from the documented run order.

提供机构:
Zenodo
创建时间:
2026-08-10
二维码
社区交流群
二维码
科研交流群
商业服务