Phenolog Identification Datasets and Supplemental Files
收藏资源简介:
Datasets generated for testing the utility of using computational techniques to assess phenotype similarity and identify phenologs. Phenologs are defined as similar phenotypes with hypothesized shared genetic basis. Identifying phenotypes mentioned in literature and predicting their similarity to other phenotypes enables candidate gene prediction in order to discover new genotype to phenotype relationships. In order to calculate the similarity between two phenotype descriptions, the descriptions first need to be converted into a computable format. Computable representations of phenotypes include EQ statements comprised of ontology terms, or embeddings into numerical vectors. These datasets correspond to the code available here. Send any feedback, questions, or suggestions to irbraun at iastate dot edu. Files: annotations.zip - Includes results of semantic annotation using NOBLE Coder, NCBO Annotator and a Naïve Bayes Classifier to map ontology terms to phenotype descriptions, as well as EQ statements annotations combining multiple ontology terms. similarity_networks.zip - Network files where edge values are given for the similarity between phenotype as calculated by each method. learned_thresholds.zip - Average similarity threshold values learned for classifying each Arabidopsis gene into a functional category. classification_matrices.zip - Matrices specifying the predicted functional category for each Arabidopsis gene. classification_summaries.xlsx - Spreadsheets summarizing the functional categorization results using each method.



