Diamond formatted protein database for taxonomic classification
收藏资源简介:
This is a diamond formatted database (diamond version 0.9.22) built on December 14th 2018. The database contains a total of 17,694,143 sequences: 2,708,401 protein sequences from the taxmapper database (commit 450d337), containing 121 unique taxa. 14,976,193 protein sequences from 1055 unique fungal taxa (downloaded from JGI 1000 fungi project on November 23 2018). 9,549 protein sequences from the <em>Hygrophorus russula</em> genome obtained from Genbank (accession GCA_003314125.1) on November 28 2018. The <em>Hygrophorus russula</em> protein sequences were obtained by running Augustus (v. 3.2.3) gene caller on the genomic fasta file using the laccaria_bicolor gene model. Taxonomic information was built into the diamond database by running: <pre><code class="language-bash">zcat fasta.gz | diamond makedb -d diamond -p 4 --taxonmap taxonmap.gz --taxonnodes nodes.dmp</code></pre> The nodes.dmp file was obtained from the taxdump.tar.gz file on December 11 2018.



