Data (part 1) from Illuminating the functional landscape of the dark proteome across the Animal Tree of Life through natural language processing models
收藏资源简介:
Part 1 contains: longest_isoforms_gopredsim_seqvec_A_to_E.tar.gz : GOPredSim GO annotation using SeqVec model for the longest isoform or species whose code first letter goes from A to E. longest_isoforms_deepgoplus.tar.gz : DeepGOPlus GO annotation for the longest isoform of all species. all_isoforms_deepgoplus.tar.gz : DeepGOPlus GO annotation for all isoforms for a subset of 102 species. all_isoforms_eggnog_filtered_goterms.tar.gz : eggNOG-mapper functional annotation for all isoforms for a subset of 102 species. annotated_genes_eggnog.txt : list of genes that have a GO term annotation from eggNOG-mapper. go_per_gene_filters_stats_go_cat_subset.tar.gz : number of GO terms per gene for all filters applied and all functional annotation methods for a subset of 102 species. semsim_longest_eggnog_seqvec_go_category.tar.gz : semantic similarity separated by GO category between the GO annotations from eggNOG-mapper and GOPredSim (SeqVec model) for the longest isoform of a subset of 102 species. semsim_longest_eggnog_prott5_go_category.tar.gz : semantic similarity separated by GO category between the GO annotations from eggNOG-mapper and GOPredSim (ProtT5 model) for the longest isoform of a subset of 102 species. semsim_longest_eggnog_deepgoplus_go_category.tar.gz : semantic similarity separated by GO category between the GO annotations from eggNOG-mapper and DeepGOPlus for the longest isoform of a subset of 102 species. isoforms_stats.tar.gz : number of genes with functional annotation when using all isoforms or just the longest of a subset of 102 species for all methods used. semsim_allvslong_go_cat.txt : semantic similarity separated by GO category between the GO annotations when using all isoforms or just the longest of a subset of 102 species. taxonomy_lookup_dataset.tar.gz : Uniprot taxonomy information used for obtaining the taxonomic origin of the proteins in the lookup dataset and results. CSCU1_unannotated_toxins.txt : list of proteins from Centruroides sculpturatus predicted to have toxin activity but not being annotated by eggNOG-mapper. ToxinPred2_0_output_CSCU_putative_toxins.pdf : ToxinPred2 output for the prediction of toxin activity of the proteins in the CSCU1_unannotated_toxins.txt file. CSCU1_DN0_c0_g23215_i1p1_20900.result.zip : ColabFold predicted protein structure of a previously unannotated putative tachylectin-5 homolog in C. sculpturatus. structures_2024-2-19-14-15-54.zip : PDB pairwise structure alignment result for C. sculpturatus new and previously identified tachylectin-5 with Tachypleus tridentatus tachylectin-5A (PDB 1JC9). XP_023224331_3bbdb.result.zip : ColabFold predicted protein structure of previously tachylectin-5-like XP_023224331 protein in C. sculpturatus. CG11373_92f17.result.zip : ColabFold predicted protein structure of CG11373 in Drosophila melanogaster. all_species_topgo_input.tar.gz : topGO input GOPredSim-ProtT5 GO annotations for all species. GO_enrichment_results_per_phylum.tar.gz : per-phylum combined results of per-species GO enrichment results. GO_enrichment_results_per_species.tar.gz : GO enrichment of GO terms of genes not annotated by eggNOG-mapper for all species. cluster_unassigned_agr.txt : manual assignment of GO cluster representatives that could not be assigned to a Alliance for Genomics Resources (AGR) GO Slim GO category. cluster_unassigned_agr_not_sure.txt : explanation of the manual assignment of GO cluster representatives that could not be assigned to a AGR GO Slim GO category. common_GOs_BP_all_phyla.txt : Biological process (BP) GO terms that were commonly enriched in all animal phyla. common_GOs_CC_all_phyla.txt : Cellular component (CC) GO terms that were commonly enriched in all animal phyla. common_GOs_MF_all_phyla.txt : Molecular function (MF) GO terms that were commonly enriched in all animal phyla. GO_clusters_comparison_per_phylum.txt : number of clusters predicted by several clustering methods for each GO category for each of the animal phyla. GOslim_agr_unassigned_GOs_clustersBP.txt : GO cluster representatives of clustered BP GO terms that could not be assigned to a Alliance for Genomics Resources (AGR) GO Slim GO category. GOslim_agr_unassigned_GOs_clustersCC.txt : GO cluster representatives of clustered CC GO terms that could not be assigned to a Alliance for Genomics Resources (AGR) GO Slim GO category. GOslim_agr_unassigned_GOs_clustersMF.txt : GO cluster representatives of clustered MF GO terms that could not be assigned to a Alliance for Genomics Resources (AGR) GO Slim GO category. GOslim_unassigned_GO_cats_proportion_per_phylum.txt : proportion of unassigned GO terms for each of the GO terms categories for 2 GO Slims (AGR and general) per phylum. unique_GOs_BP_all_phyla.txt : Biological process (BP) GO terms that were exclusively enriched in one animal phyla. unique_GOs_CC_all_phyla.txt : Cellular component (CC) GO terms that were exclusively enriched in one animal phyla. unique_GOs_MF_all_phyla.txt : Molecular function (MF) GO terms that were exclusively enriched in one animal phyla. genes_after_filters.tar.gz : genes that remain after the different filters for each of the methods. disorder_all.txt : disorder prediction of all proteins. Taxon_list_subset.tsv : species metadata for plotting. metazoa_phyla_tree.nwk : animal phyla phylogenetic tree used for plotting. subset_all_isoforms_tree.nwk : subset of 93 animal species phylogenetic tree used for plotting. Fantasia_computer_resources.txt : computer resources used by FANTASIA. model_organisms_tree.nwk : model organisms phylogenetic tree used for plotting.



