kaiju_mycobacterium
收藏资源简介:
Kaiju database – Mycobacterium subset (2024 release) This dataset provides a custom Kaiju database containing only protein sequences from the genus Mycobacterium, extracted from the NCBI NR/RefSeq repositories (August 2024).The database was built to optimize the taxonomic classification of sequencing reads from Mycobacterium tuberculosis and related species, significantly reducing computational requirements compared to the full Kaiju NR database (~100 GB). This subset includes representative genomes from Mycobacterium tuberculosis, M. bovis, M. africanum, M. smegmatis, and other clinically or environmentally relevant species within the genus. Contents: kaiju_db_mycobacterium_2024.fmi — Kaiju formatted database index nodes.dmp, names.dmp — NCBI taxonomy mapping files Total size: ~20 GBKaiju version: compatible with ≥ 1.9.0Reference source: NCBI NR/RefSeq (retrieved August 2024) Use case:Designed for pipelines performing taxonomic classification and contamination screening of Mycobacterium sequencing data, enabling faster execution while maintaining taxonomic resolution at the species level. Recommended citation: Kaiju database – Mycobacterium subset (2024 release). Zenodo. https://10.5281/zenodo.17554952 Menzel, P., Ng, K. L., & Krogh, A. (2016). Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nature Communications, 7, 11257. https://doi.org/10.1038/ncomms11257



