PheKnowLator Human Disease Knowledge Graphs - Build Data (Processed)
收藏资源简介:
<strong>RELEASE V2.1.0 KNOWLEDGE GRAPH: PROCESSED DATA SOURCES </strong> <strong>Release:</strong> v2.1.0 The goal of this build was to create a knowledge graph that represented human disease mechanisms and included the central dogma. The data sources utilized in this release include many of the sources used in the initial release, as well as some new data made available by the Comparative Toxicogenomics Database and experimental data from the Human Protein Atlas. Data sources are listed by type (Ontology and Data not represented in an ontology [Database Sources]). Additional details are provided for each data source below. Please see documentation on the primary release (https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources) for additional details on each data source as well as citation information. <strong>Data Access:</strong> https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021 <strong>ONTOLOGIES</strong> Cell Ontology Cell Line Ontology Chemical Entities of Biological Interest (ChEBI) Ontology Gene Ontology Human Phenotype Ontology Mondo Disease Ontology Pathway Ontology Protein Ontology Relations Ontology Sequence Ontology Uber-Anatomy Ontology Vaccine Ontology <strong>Cell Ontology (CL)</strong> <strong>Homepage:</strong> <strong><code>GitHub</code></strong><br> <strong>Citation:</strong> Bard J, Rhee SY, Ashburner M. An ontology for cell types. Genome Biology. 2005;6(2):R21 <strong>Usage:</strong> Utilized to connect <code>transcripts</code> and <code>proteins</code> to <code>cells</code>. Additionally, the edges between this ontology and its dependencies are utilized: <strong><code>ChEBI</code></strong> <strong><code>GO</code></strong> <strong><code>PATO</code></strong> <strong><code>PRO</code></strong> <strong><code>RO</code></strong> <strong><code>UBERON</code></strong> <strong>Cell Line Ontology (CLO)</strong> <strong>Homepage:</strong> <strong><code>http://www.clo-ontology.org/</code></strong><br> <strong>Citation:</strong> Sarntivijai S, Lin Y, Xiang Z, Meehan TF, Diehl AD, Vempati UD, Schürer SC, Pang C, Malone J, Parkinson H, Liu Y. CLO: the cell line ontology. Journal of Biomedical Semantics. 2014;5(1):37 <strong>Usage:</strong> Utilized this ontology to map <code>cell lines</code> to <code>transcripts</code> and <code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized: <strong><code>CL</code></strong> <strong><code>DOID</code></strong> <strong><code>NCBITaxon</code></strong> <strong><code>UBERON</code></strong> <strong>Chemical Entities of Biological Interest (ChEBI)</strong> <strong>Homepage:</strong> <strong><code>https://www.ebi.ac.uk/chebi/</code></strong><br> <strong>Citation:</strong> Hastings J, Owen G, Dekker A, Ennis M, Kale N, Muthukrishnan V, Turner S, Swainston N, Mendes P, Steinbeck C. ChEBI in 2016: Improved services and an expanding collection of metabolites. Nucleic Acids Research. 2015;44(D1):D1214-9 <strong>Usage:</strong> Utilized to connect <code>chemicals</code> to <code>complexes</code>, <code>diseases</code>, <code>genes</code>, <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>pathways</code>, <code>phenotypes</code>, <code>reactions</code>, and <code>transcripts</code>. <strong>Gene Ontology (GO)</strong> <strong>Homepage:</strong> <strong><code>http://geneontology.org/</code></strong><br> <strong>Citations:</strong> Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA. Gene ontology: tool for the unification of biology. Nature Genetics. 2000;25(1):25 The Gene Ontology Consortium. The Gene Ontology Resource: 20 years and still GOing strong. Nucleic Acids Research. 2018;47(D1):D330-8 <strong>Usage:</strong> Utilized to connect <code>biological processes</code>, <code>cellular components</code>, and <code>molecular functions</code> to <code>chemicals</code>, <code>pathways</code>, and <code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized: <strong><code>CL</code></strong> <strong><code>NCBITaxon</code></strong> <strong><code>RO</code></strong> <strong><code>UBERON</code></strong> <strong>Other Gene Ontology Data Used:</strong> <code>goa_human.gaf.gz</code> <strong>Human Phenotype Ontology (HPO)</strong> <strong>Homepage:</strong> <strong><code>https://hpo.jax.org/</code></strong><br> <strong>Citation:</strong> Köhler S, Carmody L, Vasilevsky N, Jacobsen JO, Danis D, Gourdine JP, Gargano M, Harris NL, Matentzoglu N, McMurry JA, Osumi-Sutherland D. Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources. Nucleic Acids Research. 2018;47(D1):D1018-27 <strong>Usage:</strong> Utilized to connect <code>phenotypes</code> to <code>chemicals</code>, <code>diseases</code>, <code>genes</code>, and <code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized: <strong><code>CL</code></strong> <strong><code>ChEBI</code></strong> <strong><code>GO</code></strong> <strong><code>UBERON</code></strong> <strong>Files</strong> Other Human Phenotype Ontology Data Used: <code>phenotype.hpoa</code> <strong>Mondo Disease Ontology (Mondo)</strong> <strong>Homepage:</strong> <strong><code>https://mondo.monarchinitiative.org/</code></strong><br> <strong>Citation:</strong> Mungall CJ, McMurry JA, Köhler S, Balhoff JP, Borromeo C, Brush M, Carbon S, Conlin T, Dunn N, Engelstad M, Foster E. The Monarch Initiative: an integrative data and analytic platform connecting phenotypes to genotypes across species. Nucleic Acids Research. 2017;45(D1):D712-22 <strong>Usage:</strong> Utilized to connect <code>diseases</code> to <code>chemicals</code>, <code>phenotypes</code>, <code>genes</code>, and <code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized: <strong><code>CL</code></strong> <strong><code>NCBITaxon</code></strong> <strong><code>GO</code></strong> <strong><code>HPO</code></strong> <strong><code>UBERON</code></strong> <strong>Pathway Ontology (PW)</strong> <strong>Homepage:</strong> <strong><code>rgd.mcw.edu</code></strong><br> <strong>Citation:</strong> Petri V, Jayaraman P, Tutaj M, Hayman GT, Smith JR, De Pons J, Laulederkind SJ, Lowry TF, Nigam R, Wang SJ, Shimoyama M. The pathway ontology–updates and applications. Journal of Biomedical Semantics. 2014;5(1):7. <strong>Usage:</strong> Utilized to connect <code>pathways</code> to <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>Reactome pathways</code>. Several steps are taken in order to connect <code>Pathway Ontology</code> identifiers to <code>Reactome</code> pathways and <code>GO biological processes</code>. To connect <code>Pathway Ontology</code> identifiers to <code>Reactome</code> pathways, we use ComPath Pathway Database Mappings developed by Daniel Domingo-Fernández (PMID:30564458). <strong>Files</strong> Downloaded Mapping Data <code>curated_mappings.txt</code> <code>kegg_reactome.csv</code> Generated Mapping Data <code>REACTOME_PW_GO_MAPPINGS.txt</code> <strong>Protein Ontology (PRO)</strong> <strong>Homepage:</strong> <strong><code>https://proconsortium.org/</code></strong><br> <strong>Citation:</strong> Natale DA, Arighi CN, Barker WC, Blake JA, Bult CJ, Caudy M, Drabkin HJ, D’Eustachio P, Evsikov AV, Huang H, Nchoutmboube J. The Protein Ontology: a structured representation of protein forms and complexes. Nucleic Acids Research. 2010;39(suppl_1):D539-45 <strong>Usage:</strong> Utilized to connect <code>proteins</code> to <code>chemicals</code>, <code>genes</code>, <code>anatomy</code>, <code>catalysts</code>, <code>cell lines</code>, <code>cofactors</code>, <code>complexes</code>, <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>pathways</code>, <code>proteins</code>, <code>reactions</code>, and <code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized: <strong><code>ChEBI</code></strong> <strong><code>DOID</code></strong> <strong><code>GO</code></strong> <strong>Notes:</strong> A partial, human-only version of this ontology was used. Details on how this version of the ontology was generated can be found under the Protein Ontology section of the <code>Data_Preparation.ipynb</code> Jupyter Notebook. <strong>Files</strong> Generated Human Version Protein Ontology (PRO) <code>human_pro.owl</code> (closed with hermit reasoner) Other PRO Data Used: <code>promapping.txt</code> Generated Mapping Data Merged Gene, RNA, Protein Map: <code>Merged_gene_rna_protein_identifiers.pkl</code> Ensembl Transcript-PRO Identifier Mapping: <code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code> Entrez Gene-PRO Identifier Mapping: <code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code> UniProt Accession-PRO Identifier Mapping: <code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code> STRING-PRO Identifier Mapping: <code>STRING_PRO_ONTOLOGY_MAP.txt</code> <strong>Relations Ontology (RO)</strong> <strong>Homepage:</strong> <strong><code>GitHub</code></strong><br> <strong>Citation:</strong> Smith B, Ceusters W, Klagges B, Köhler J, Kumar A, Lomax J, Mungall C, Neuhaus F, Rector AL, Rosse C. Relations in biomedical ontologies. Genome Biology. 2005;6(5):R46. <strong>Usage:</strong> Utilizing this ontology to connect all data sources in knowledge graph. Additionally, the ontology is queried prior to building the knowledge graph to identify all relations, their inverse properties, and their labels. <strong>Files</strong> Generated RO Data <code>INVERSE_RELATIONS.txt</code> <code>RELATIONS_LABELS.txt</code> <strong>Sequence Ontology (SO)</strong> <strong>Homepage:</strong> <strong><code>GitHub</code></strong><br> <strong>Citation:</strong> Eilbeck K, Lewis SE, Mungall CJ, Yandell M, Stein L, Durbin R, Ashburner M. The Sequence Ontology: a tool for the unification of genome annotations. Genome Biology. 2005;6(5):R44 <strong>Usage:</strong> Utilized to connect <code>transcripts</code> and other genomic material like <code>genes</code> and <code>variants</code>. <strong>Files</strong> Generated Mapping Data <code>genomic_sequence_ontology_mappings.xlsx</code> <code>SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt</code> <strong>Uber-Anatomy Ontology (Uberon)</strong> <strong>Homepage:</strong> <strong><code>GitHub</code></strong><br> <strong>Citation:</strong> Mungall CJ, Torniai C, Gkoutos GV, Lewis SE, Haendel MA. Uberon, an integrative multi-species anatomy ontology. Genome Biology. 2012;13(1):R5 <strong>Usage:</strong> Utilized to connect <code>tissues</code>, <code>fluids</code>, and <code>cells</code> to <code>proteins</code> and <code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized: <strong><code>ChEBI</code></strong> <strong><code>CL</code></strong> <strong><code>GO</code></strong> <strong><code>PRO</code></strong> <strong>Vaccine Ontology (VO)</strong> <strong>Homepage:</strong> <strong><code>http://www.violinet.org/vaccineontology/</code></strong><br> <strong>Citations:</strong> He Y, Racz R, Sayers S, Lin Y, Todd T, Hur J, Li X, Patel M, Zhao B, Chung M, Ostrow J. Updates on the web-based VIOLIN vaccine database and analysis system. Nucleic Acids Research. 2013;42(D1):D1124-32 Xiang Z, Todd T, Ku KP, Kovacic BL, Larson CB, Chen F, Hodges AP, Tian Y, Olenzek EA, Zhao B, Colby LA. VIOLIN: vaccine investigation and online information network. Nucleic Acids Research. 2007;36(suppl_1):D923-8 <strong>Usage:</strong> Utilized the edges between this ontology and its dependencies: <strong><code>ChEBI</code></strong> <strong></strong>



