遇见数据集

Genomics NCD-gzip database

收藏
Zenodo2025-06-10 更新2026-05-26 收录
官方服务:

资源简介:

This is a basic database consisting of 4634 genomes sampled from the RefSeq database using the Woltka pipelin. It is intended for testing and evaluation of metagenomic classification tools. Contents The genome sequences are provided in a compressed archive (genomes.tar.gz). When unpacked, the folder structure is organized by NCBI Taxonomy ID (taxid), like so: genomes/├── taxid1/│ ├── genome1_0.fna│ └── genome1_1.fna├── taxid2/│ └── genome2_0.fna├── taxid3/│ ├── genome3_0.fna│ ├── genome3_1.fna│ └── genome3_2.fna Each top-level directory corresponds to a taxonomic ID and contains one or more genome FASTA files in .fna format. Additional Files fold1_list.txt and fold1_testing_list.txt: Lists of genome TaxIDs used for training and testing, respectively. These are included to support reproducible benchmarking of metagenomic classifiers.

提供机构:
Zenodo
创建时间:
2025-06-10
二维码
社区交流群
二维码
科研交流群
商业服务