MNBC: a multithreaded Minimizer-based Naïve Bayes Classifier for improved metagenomic sequence classification
收藏资源简介:
These files provide supplementary data underlying the article (see Figure 1 in the article): 37345_filtered_training_and_test_genomes_list.txt: Refseq assembly sequence filenames of the 37345 filtered training and test genomes taxonomy_37345_filtered_training_and_test_genomes.txt: Taxonomy file for all 37345 filtered training and test genomes Uniform_reference_database_31991_training_genomes_assemblyID_list.txt: Refseq assembly IDs of the 31991 training genomes in the uniform reference database uniform_reference_database.tar.gz_1 to uniform_reference_database.tar.gz_10: Merge them into a single file using the cat command. The folder produced by decompressing this file is the uniform reference database. taxonomy_uniform_reference_database.txt: Taxonomy file for the uniform reference database (i.e. the 31991 training genomes) 4964_prokaryotic_test_genomes_assemblyID_list.txt: Refseq assembly IDs of the 4964 prokaryotic test genomes 390_viral_test_genomes_assemblyID_list.txt: Refseq assembly IDs of the 390 viral test genomes testLongReads_prok_C0.106.fasta.gz: 2089493 1Kb-long prokaryotic test reads randomly generated from the 4964 prokaryotic test genomes (0.106 coverage, in FASTA format) testLongReads_virus_C6.fasta.gz: 100983 1Kb-long viral test reads randomly generated from the 390 viral test genomes (6 coverage, in FASTA format) CAMI2_reference_database_16864_genomes_list.txt: Refseq assembly sequence filenames of the 16864 genomes and chromosomes in the reference database for CAMI2 taxonomy_CAMI2_reference_database.txt: Taxonomy file for the CAMI2 reference database



