Phold Manuscript Supplementary Files Too Large For Github
收藏资源简介:
Contains all supplementary files (mostly Foldseek search and clustering outputs) that go with the `phold-analysis` repository (https://github.com/gbouras13/phold-analysis) that are too large for GitHub. `afdb50_phold_bfvd_nomburg.tar.gz` - Foldseek clustering results from Phold, BFVD, Nomburg and AFDB databases `efam_res_3_iterations.m8.gz` - Foldseek search results of efam structures vs PHROG structures (used to assign PHROG groups to efam proteins) `envhog_mmseqs_consensus.fasta.gz` - enVhog consensus protein sequences (FASTA) `envhog_res_3_iterations.m8.gz` - Foldseek search results of enVhog structures vs PHROG structures (used to assign PHROG groups to efam proteins) `all_phold_structures_plddt_70.fasta.gz` - all amino acid protein sequences used to train the LoRA models `all_phold_structures_plddt_70_ss.fasta.gz` - all 3Di protein sequences used to train the LoRA models `BFVD.m8.gz` - raw Foldseek search output of all Phold DB 3.16M vs BFVD proteins. NOTE - `AFDB.m8.gz` raw Foldseek search of all Phold DB 3.16M vs AFDB proteins is 124GB compressed and too large for zenodo. Please contact me george.bouras@adelaide.edu.au for more information `phrog_phrog_3_res_iterations.m8.gz` - raw Foldseek search output of all "unknown function" PHROG proteins (v4 release) against all PHROG proteins with non-unknown function `singleton_phrog_res_3_iterations.m8.gz` - raw Foldseek search output fo all PHROG singletons vs other PHROG protein structures (i.e. all in PHROGs 1-38880)



