PheKnowLator Human Disease Knowledge Graph Benchmarks Archive
收藏资源简介:
<strong>PKT-KG KG Benchmark Builds</strong> <strong>Websites</strong> https://github.com/callahantiff/PheKnowLator/wiki/Benchmarks-and-Builds https://github.com/callahantiff/PheKnowLator/wiki/Archived-Builds <strong>Preprint: </strong>http://arxiv.org/abs/2307.05727 The PKT Human Disease KG was built to model mechanisms of human disease, which includes the Central Dogma and represents multiple biological scales of organization including molecular, cellular, tissue, and organ. The knowledge representation was designed in collaboration with a PhD-level molecular biologist (Figure). <strong>PKT Human Disease KG Benchmark Build Information</strong> To enable customization in the way that knowledge is represented when constructing a KG, three configurable parameters are provided: <strong>Knowledge Model. </strong>Following Semantic Web standards, PKT-KG defines a KG as K=〈T, A〉, where T is the TBox and A is the ABox. The TBox represents the taxonomy of a particular domain. It describes classes, properties/relationships, and assertions that are assumed to generally hold within a domain (e.g., a gene is a heritable unit of DNA located in the nucleus of cells [Figure 7a]). The ABox describes attributes and roles of instances of classes (i.e., individuals) and assertions about their membership in classes within the TBox (e.g., A2M is a type of gene that may cause Alzheimer’s Disease [Figure 7b]). PKT KGs are logically grounded in one or more OBO Foundry ontology. Database entities (i.e., entities from a data source that is not an OBO Foundry ontology) are added to the core OBO Foundry ontologies using either a TBox (i.e., class-based) or ABox (i.e., instance-based) knowledge model. For the class-based approach, each database entity is made a subclass of an existing core OBO Foundry ontology class (see the “Class-based” section of Supplementary Table 14). For the instance-based approach, each database entity is made an instance of an existing core OBO Foundry ontology class (see the “Instance-based” section of Supplementary Table 14). Both approaches require the alignment of database entities to an existing core OBO Foundry ontology class, which is managed by a dictionary that is constructed using tools in the Process Data Element of the Knowledge Graph Construction Resources component . <strong>Relation Strategy. </strong>PKT-KG provides two relation strategies. The first strategy is standard or directed relations, through a single directed edge (e.g., “gene causes phenotype”). The second strategy is inverse or bidirectional relations, through inference if the relation is from an ontology like the RO (e.g., “chemical participates in pathway” and “pathway has participant chemical”) or through inferring implicitly symmetric relations for edge types that represent biological interactions (e.g., gene-gene interactions). <strong>Semantic Abstraction. </strong>KGs built using expressive languages like OWL are structurally complex and composed of triples or edges that are logically necessary but not biologically meaningful (e.g., anonymous subclasses used to express TBox assertions with all-some quantification). PKT-KG currently uses the OWL-NETS (PMC5737627)semantic abstraction algorithm to convert or transform complex KGs into hybrid KGs. OWL-NETS v2.0 includes additional functionality that harmonizes a semantically abstracted KG to be consistent with a class- or instance-based knowledge model. For class-based knowledge models, all triples containing rdf:type are updated to rdfs:subClassOf and for instance-based knowledge models, all triples containing rdfs:subClassOf are updated to rdf:type. For additional details, see OWL-NETS v2.0 documentation. The PKT Human Disease KG was constructed using 12 OBO Foundry ontologies, 31 Linked Open Data sets, and results from two large-scale experiments (Supplementary Table 12). The 12 OBO Foundry ontologies were selected to represent chemicals and vaccines (i.e., ChEBI and Vaccine Ontology [VO]), cells and cell lines (i.e., Cell Ontology [CL], Cell Line Ontology [CLO]), gene/gene product attributes (i.e., Gene Ontology [GO]), phenotypes and diseases (i.e., Human Phenotype Ontology [HPO], Mondo Disease Ontology [Mondo]), proteins, including complexes and isoforms (i.e., PRO), pathways (i.e., Pathway Ontology [PW]), types and attributes of biological sequences (i.e., SO), and anatomical entities (Uberon). The RO is used to provide relationships between the core OBO Foundry ontologies and database entities. Tthe PKT Human Disease KG contained 18 node types and 33 edge types. Note that the number of nodes and edge types reflects those that are explicitly added to the core set of OBO Foundry ontologies and does not take into account the node and edge types provided by the ontologies. These nodes and edge types were used to construct 12 different PKT Human Disease benchmark KGs by altering the Knowledge Model (i.e., class- vs. instance-based), Relation Strategy (i.e., standard vs. inverse relations), and Semantic Abstraction (i.e., OWL-NETS (yes/no) with and without Knowledge Model harmonization [OWL-NETS Only vs. OWL-NETS + Harmonization]) parameters. Benchmarks within the PheKnowLator ecosystem are different versions of a KG that can be built under alternative knowledge models, relation strategies, and with or without semantic abstraction. They provide users with the ability to evaluate different modeling decisions (based on the prior mentioned parameters) and to examine the impact of these decisions on different downstream tasks. <strong>Files in this Directory</strong> <strong>PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</strong><em>: </em>contains information on the different types of files that are output for each build. <strong>PKT-KG Benchmark Build - Full Data Access</strong> All Builds are made 100% publicly available and can be accessed in a number of different ways. Please note that both of the ways listed below provides access to all data used to generate the KGs and all generated KG files. <strong>Google Cloud Storage Bucket: </strong>https://console.cloud.google.com/storage/browser/pheknowlator <strong>GitHub Wiki:</strong> https://github.com/callahantiff/PheKnowLator/wiki/Archived-Builds <strong>PKT-KG Benchmark Build - KG Data Access</strong> <em><strong>v1.0.0</strong></em> <em>KGs</em>: https://zenodo.org/record/7030201 <em>Embeddings:</em> https://zenodo.org/record/7030189 <em><strong>All Other Versions</strong></em> Build Type v2.0.0 v2.1.0 v3.0.2 Class-based + StandardRelations + OWL MAY2020 JAN2021 FEB2021 MAY2021 JUN2021 JUL2021 AUG2021 SEP2021 OCT2021 NOV2021 Class-based + StandardRelations + OWL-NETS MAY2020 JAN2021 FEB2021 MAY2021 JUN2021 JUL2021 AUG2021 SEP2021 OCT2021 NOV2021 Class-based + InverseRelations + OWL MAY2020 JAN2021 FEB2021 MAY2021 JUN2021 JUL2021 AUG2021 SEP2021 OCT2021 NOV2021 Class-based + InverseRelations + OWL-NETS MAY2020 JAN2021 FEB2021 MAY2021 JUN2021 JUL2021 AUG2021 SEP2021 OCT2021 NOV2021 Instance-based + StandardRelations + OWL MAY2020 JAN2021 FEB2021 MAY2021 JUN2021 JUL2021 AUG2021 SEP2021 OCT2021 NOV2021 Instance-based + StandardRelations + OWL-NETS MAY2020 JAN2021 FEB2021 MAY2021 JUN2021 JUL2021 AUG2021 SEP2021 OCT2021 NOV2021 Instance-based + InverseRelations + OWL MAY2020 JAN2021 FEB2021 MAY2021 JUN2021 JUL2021 AUG2021 SEP2021 OCT2021 NOV2021 Instance-based + InverseRelations + OWL-NETS MAY2020 JAN2021 FEB2021 MAY2021 JUN2021 JUL2021 AUG2021 SEP2021 OCT2021 NOV2021



