遇见数据集

LipoKG: A Comprehensive Knowledge Graph Dataset for Lipoprotein Metabolism Research

收藏
Zenodo2026-10-01 更新2026-10-01 收录
官方服务:

资源简介:

LipoKG is a FAIR-compliant knowledge graph integrating multi-source data for lipoprotein metabolism research. It combines protein-protein interactions (STRING v12.0), pathogenic variants (ClinVar), disease and trait associations (DisGeNET, OMIM, GLGC 2021, Orphanet and expert curation) and pathway information (KEGG, Reactome, WikiPathways) into a single graph in which the lipoprotein particles - chylomicrons, VLDL, IDL, LDL, HDL, Lp(a) and remnants - are modelled as first-class entities. The graph contains 6,214 core nodes and 36,568 core edges across 7 node and 7 relationship types; the extended Neo4j schema declares a further 425 nodes and 4,929 relationships, giving 6,639 nodes and 41,497 relationships in total across 32 node types and 26 relationship types. Pathway completeness across the three construction databases is 74.9% (259 of 346 reference genes): KEGG hsa05417 72.7% (157/216), Reactome plasma-lipoprotein pathways 85.9% (61/71) and WikiPathways lipid pathways 79.3% (73/92). Four independent benchmarks assess coverage beyond the construction sources: GO:0042157 "lipoprotein metabolic process" 92.3% (132/143 genes); GLGC 2021 genome-wide significant lipid loci 64.9% (244/376); ClinGen lipid-related dosage-sensitive genes 100% (25/25); and an expert-curated 54-gene clinical reference set 100% (54/54). The dataset supports systematic annotation-gap discovery, target prioritisation, drug repurposing and precision-medicine applications in cardiovascular disease. Version 1.2.1 re-retrieves the KEGG layer as hsa05417 ("Lipid and atherosclerosis", 216 genes) and the Reactome and WikiPathways reference sets, and adds data/LICENSE.md - a per-file and per-source layered licence. Components whose sources permit it are released under CC0; layers derived from more restrictive sources retain those terms, and the repository-level CC BY 4.0 label does not override them. The dataset includes the six extended-schema CSV exports under data/extended/, a per-file inventory with byte sizes and row counts (Table S25), and a gene-level annotation coverage matrix (Table S24): of the 1,852 proteins in the interaction graph, 252 (13.6%) carry at least one annotation layer beyond STRING connectivity and 1,600 (86.4%) are represented by STRING edges only.

提供机构:
Zenodo
创建时间:
2026-10-01
二维码
社区交流群
二维码
科研交流群
商业服务