遇见数据集

kaiju_mycobacterium

收藏
Zenodo2025-11-07 更新2026-05-26 收录
官方服务:

资源简介:

Kaiju database – Mycobacterium subset (2024 release) This dataset provides a custom Kaiju database containing only protein sequences from the genus Mycobacterium, extracted from the NCBI NR/RefSeq repositories (August 2024).The database was built to optimize the taxonomic classification of sequencing reads from Mycobacterium tuberculosis and related species, significantly reducing computational requirements compared to the full Kaiju NR database (~100 GB). This subset includes representative genomes from Mycobacterium tuberculosis, M. bovis, M. africanum, M. smegmatis, and other clinically or environmentally relevant species within the genus. Contents: kaiju_db_mycobacterium_2024.fmi — Kaiju formatted database index nodes.dmp, names.dmp — NCBI taxonomy mapping files Total size: ~20 GBKaiju version: compatible with ≥ 1.9.0Reference source: NCBI NR/RefSeq (retrieved August 2024) Use case:Designed for pipelines performing taxonomic classification and contamination screening of Mycobacterium sequencing data, enabling faster execution while maintaining taxonomic resolution at the species level. Recommended citation: Kaiju database – Mycobacterium subset (2024 release). Zenodo. https://10.5281/zenodo.17554952 Menzel, P., Ng, K. L., & Krogh, A. (2016). Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nature Communications, 7, 11257. https://doi.org/10.1038/ncomms11257

提供机构:
Zenodo
创建时间:
2025-11-07
二维码
社区交流群
二维码
科研交流群
商业服务