遇见数据集

Pharmacogenomic Knowledge Graph (PGx-KG) Dataset

收藏
Zenodo2025-09-26 更新2026-05-26 收录
官方服务:

资源简介:

Pharmacogenomic Knowledge Graph (PGx-KG) Dataset This repository contains the processed pharmacogenomic knowledge graph (PGx-KG) prepared for benchmarking link prediction models. The graph integrates curated relationships from PharmGKB, ClinVar, SIDER, and Reactome to standard biomedical vocabularies so that entities and relations align across resources. Contents nodes.csv: Consolidated catalogue of biomedical entities with harmonized identifiers, semantic types, and source provenance. edges.csv: Full edge list covering all relation types with integer IDs referencing the node maps. triples/: Relation-specific subject–predicate–object triples in CSV format, named as <REL>__HEADTYPE__TAILTYPE. relations/: Per-relation tables with descriptive metadata, confidence scores, and source annotations. id_maps/: Mapping tables that link integer IDs to canonical biomedical identifiers (HGNC, RxNorm, MeSH, ChEBI, etc.). degree/: Per-relation in/out degree distributions for downstream analysis or sampling. splits/: Stratified train/validation/test partitions for each relation type, with only CSV versions (including reproducible random/ test/train/val). stats.json: Summary statistics computed after harmonization (entity counts, relation counts,). Data Statistics Total nodes: 3,744,727 Total edges: 9,645,367 Node types: Variants, Genes, Diseases, Drugs, ADRs, Pathways Relation types: VAR_ASSOC_DIS, VAR_IN_GENE, DRUG_CAUSES_ADR, GENE_IN_PATH, DRUG_IN_PATH, GENE_AFFECTS_DRUG How to Use Load the relation-specific triples in triples/ (or the global edge list in edges.*) into your preferred graph ML framework (e.g., PyTorch Geometric, DGL). Use the mapping tables in id_maps/ (e.g., GENE_ids.csv, DRUG_ids.csv) to translate integer IDs back to canonical biomedical identifiers. Evaluate models on the curated train/validation/test splits under splits/<REL>/ and report metrics such as MRR, Hits@K, or AUROC. Source Databases This dataset was derived from the following resources (not redistributed here): PharmGKB (https://www.pharmgkb.org/) ClinVar (https://www.ncbi.nlm.nih.gov/clinvar/) SIDER (http://sideeffects.embl.de/) Reactome (https://reactome.org/) License This processed dataset is released under the CC BY 4.0 License. You are free to share and adapt, provided appropriate credit is given. Citation Citation If you use this dataset, please cite both the dataset and the preprint: Faruk, M.O. (2025). Pharmacogenomic Knowledge Graph (PGx-KG): Processed Dataset for Link Prediction. Zenodo. https://doi.org/10.5281/zenodo.17189995 Faruk, M.O. (2025). A large-scale pharmacogenomic knowledge graph for drug-gene-variant-disease discovery. medRxiv. https://doi.org/10.1101/2025.09.24.25336269v1

提供机构:
Zenodo
创建时间:
2025-09-24
二维码
社区交流群
二维码
科研交流群
商业服务