遇见数据集

Phenotype-Associated Variants in Saudi Arabia

收藏
Zenodo2026-06-24 更新2026-06-28 收录
官方服务:

资源简介:

Phenotype-Associated Variants in Saudi Arabia PAVS (Phenotype-Associated Variants in Saudi Arabia) is a curated genotype–phenotype database integrating rare disease case data from Saudi Arabian clinical cohorts, mixed-population cohorts, and literature-derived phenopackets. All cases are encoded as GA4GH Phenopackets v2.0 and integrated into an RDF knowledge graph using the PAVS ontology, HPO, OMIM, and related biomedical vocabularies. This archive contains: PAVS cases TSV (PAVS_cases.tsv): flat, tab-separated table of all 7,510 clinical cases with one row per case, containing case ID, cohort source, sex, consanguinity, family history, solved status, disease, gene, variant (HGVS + VCF), zygosity, ACMG classification, VEP consequence, gnomAD scores, and HPO terms. Suitable for direct use in R, Python, or Excel. PAVS phenopackets (combined and individual): 7,510 clinical cases in GA4GH Phenopackets v2.0 format, comprising 5,132 Saudi cases (3,710 from published clinical cohort studies [Alfares et al. 2017; Monies et al. 2017, 2019] and 1,422 curated from the Saudi case-report literature), 522 mixed-population clinical cases (Ziats et al.), and 1,856 DDD clinical cases. Literature phenopackets (9,588 cases from the Phenopackets Store) are included in the RDF knowledge graph (literature.ttl.gz) but are not redistributed as individual JSON files here — they remain available at https://github.com/monarch-initiative/phenopacket-store. RDF knowledge graph (gzip-compressed Turtle): six named graphs covering cases, gene records, HPO annotation data, HPO information content, literature phenopackets, and dataset metadata, totalling approximately 2.4 million RDF triples. Designed for deployment in OpenLink Virtuoso (SPARQL 1.1 endpoint) and queryable via the live instance at https://pavs.phenomebrowser.net. VoID dataset description (void.ttl): W3C VoID vocabulary metadata describing the knowledge graph, its named graphs, and 28 external linksets to reference ontologies and databases. Also served live at https://pavs.phenomebrowser.net/.well-known/void. Manually curated case data (XLSX): Source spreadsheet for the 1,422 manually curated Saudi cases (M-cohort), prior to phenotype normalization and phenopacket conversion. The PAVS web interface, API (FastAPI/OpenAPI), and all processing code are available at: https://github.com/bio-ontology-research-group/pavs The Arabic translation of HPO used in this work is archived separately at: https://doi.org/10.5281/zenodo.19311231 This dataset accompanies the manuscript: "A standardized database of phenotype-associated variants from Saudi Arabian rare disease patients", submitted to Scientific Data (Nature). Version 1.1.0 (2026-06-16) — changes from 1.0.0 Severity modifiers attach only to the phenotype term they qualify (204 instances across 189 cases, ≈0.4% of assignments; independently graded at 99% by expert review). Added pavs:rsId / rdfs:seeAlso dbSNP links to genomic variants in the cases graph (1,345 statements over 1,335 variants), restoring the dbSNP/TogoVar linkset. pavs:storeOverlap / dct:source provenance flags and the corrected cohort assignment (the mixed-population cohort is no longer flagged Saudi: 5,132 Saudi, 522 mixed, 1,856 DDD). Recovered source PMIDs for 59 manually curated Saudi case reports whose reference cell had been recorded without a PMID (54 from the Shaheen et al. 2019 congenital-microcephaly cohort, PMID 30214071, per the curated spreadsheet; the remainder from neighbouring publication blocks). Source-PMID coverage rises from 1,361/1,422 (95.7%) to 1,420/1,422 (99.9%); only two cases citing a non-PubMed journal remain without a PMID. As a consequence, two additional cases (M0001311, M0001411) now match a shared Phenopacket Store publication, so the store-overlap count rises from 129 to 131 (five shared publications, unchanged). Metadata and VoID descriptions bumped to version 1.1.0. Files in This Folder PAVS_cases.tsv — Flat TSV: one row per case, all clinical + variant + phenotype fields (~5 MB) PAVS_phenopackets.json — All 7,510 clinical phenopackets in a single JSON array (GA4GH Phenopackets v2.0) (~46 MB) PAVS_phenopackets_individual.zip — Individual JSON files for each of the 7,510 PAVS clinical cases (~14 MB) cases.ttl — RDF/Turtle: PAVS clinical case instances (named graph pavs:CasesGraph) (~5.9 MB) genes.ttl — RDF/Turtle: gene and disease records (named graph pavs:GenesGraph) (~9.6 MB) hpoa.ttl — RDF/Turtle: HPO annotation data for diseases (named graph pavs:HPOAGraph) (~60 MB) hpo_ic.ttl — RDF/Turtle: HPO term information content values (named graph pavs:ICGraph) (~2.7 MB) literature.ttl — RDF/Turtle: literature phenopackets from Phenopackets Store (named graph pavs:LiteratureGraph) (~16 MB) metadata.ttl — RDF/Turtle: dataset metadata, provenance, VoID linksets (named graph pavs:MetadataGraph) (~14 KB) void.ttl — W3C VoID dataset description (standalone copy of what is served at /.well-known/void) (~14 KB) PAVS_manually_curated_cases.xlsx — Source spreadsheet for 1,422 manually curated Saudi cases (M-cohort) (~241 KB) evaluation_negation.xlsx — Expert grading of all 44 negated HPO assignments by both curators (M.A., P.N.S.), each with the raw clinical text (~17 KB) evaluation_modifier.csv — Expert grading of a 100-assignment random sample of severity modifiers by P.N.S., each with the raw clinical text (~21 KB) Note on individual vs. combined phenopackets: PAVS_phenopackets_individual.zip contains the 7,510 PAVS clinical cases only (Saudi + Ziats + DDD). The literature phenopackets (9,588 cases) are included in PAVS_phenopackets.json and literature.ttl.gz but were sourced from the Phenopackets Store (https://github.com/monarch-initiative/phenopacket-store) and are not duplicated as individual files here.Known Limitations / Data Quality The workflow tags a small fraction of HPO assignments with a negation status (phenotype absent/excluded) or a severity modifier (e.g. mild, severe). Two domain experts (M.A., P.N.S.) independently validated these against the source clinical text. The graded evaluation tables are included in this archive (evaluation_negation.xlsx, evaluation_modifier.csv).

提供机构:
Zenodo
创建时间:
2026-06-24
二维码
社区交流群
二维码
科研交流群
商业服务