遇见数据集

Strail-level metagenomic simulated evaluation datasets and results

收藏
Zenodo2026-03-13 更新2026-05-26 收录
官方服务:

资源简介:

Badread simulated metagenomic datasets (nanopore2023 error model): sim_small (4 strains): This dataset includes four different strain references: two strains from Adlercreutzia equolifaciens and two from Streptococcus anginosus, with varying relative abundances. sim_medium (15 strains): This dataset contains 15 strain references distributed across five bacterial species (i.e., Helicobacter pylori, Cutibacterium acnes, Streptococcus intermedius, Streptococcus mutans, Lactococcus lactis), with each species represented by three strains. The strains exhibit different abundance levels at the species level while strains of one species are equally abundant. sim_expanded (30 strains): An extension of the sim_medium dataset, incorporating 15 additional strains from distinct species and maintaining variable abundance levels across species. Newly added species: Porphyromonas asaccharolytica, Isoptericola variabilis, Paraprevotella xylaniphila, Lancefieldella parvula, Alistipes shahii, Fusobacterium mortiferum, Blautia hansenii, Streptococcus parasanguinis, Clostridium saccharolyticum, Campylobacter ureolyticus, Campylobacter concisus , Bacteroides caccae , Roseburia hominis, Fusobacterium ulcerans, Acidaminococcus intestini. data_for_figures.zip archive contains the processed outputs used to generate Figures 2–6 and Supplementary Figures S4–S7 in the manuscript "MADRe: Strain-Level Metagenomic Classification Through Assembly-Driven Database Reduction". The files include read classification results, read count abundances, clustering information, and ground-truth annotations used for evaluation. These processed outputs were used to compute performance metrics and generate the figures reported in the manuscript. Directory overview fig2_3_S4Contains results used for Figures 2–3 and Supplementary Figure S4.These files correspond to simulated datasets with increasing complexity: sim_small sim_medium sim_expanded Outputs are provided for the following tools: MADRe, MADRe_RC, MORA, Centrifuger, Kraken2, AugPatho-ID, and AugPatho-REP.The directory also contains clustering files and ground-truth read assignments used for evaluation. fig4_figS5Contains results used for Figure 4 and Supplementary Figure S5.These experiments evaluate performance across different sequencing technologies and error profiles. The datasets include: sim_high sim_low_hifi sim_low_r10 sim_low_r9 fig5_figS6Contains results used for Figure 5 and Supplementary Figure S6.These experiments correspond to Zymo microbial community datasets: d6322_ont d6331_hifi d6331_ont Ground-truth read mappings are provided for these datasets. fig6_S7Contains aggregated abundance outputs used for Figure 6 and Supplementary Figure S7.These files include genome-level, cluster-level, and species-level abundance estimates for the evaluated tools. File types read_labels filesContain read-level classification results produced by the evaluated tools. read_classification.outContains read-to-reference assignments produced by MADRe. rc_abundances.outContains estimated genome abundances. rc_abundances.clusters.outContains abundances aggregated over clusters of highly similar genomes. clusters.txtLists reference genomes belonging to each cluster. representatives.txtLists representative genomes selected for each cluster. paf filesFor some datasets only PAF files containing read-to-reference mappings are provided. These mapping files were directly used to extract the classification results and compute the evaluation metrics. Kraken2 output filesFor Kraken2, the direct classification output file is provided and was used directly to extract read-level assignments and compute abundance estimates. Ground truth filesContain the expected read-to-reference assignments used to evaluate classification accuracy. Notes All files in this archive represent processed outputs used to generate the figures in the MADRe manuscript. Raw sequencing data are available from the original datasets referenced in the manuscript.

提供机构:
Zenodo
创建时间:
2025-06-27
二维码
社区交流群
二维码
科研交流群
商业服务