Cross-architecture ensembling of DNA foundation models improves the precision and stability of chimera detection in long-read metagenomic bins
收藏资源简介:
Processed embeddings and analysis tables for a chimera-detection benchmark of Evo 2 (7B) vs DNABERT-S (117M) DNA foundation models on metagenome-assembled genomes. Covers two benchmarks: (1) CAMI2 Nanopore (131 MAGs, 32 true chimeras) — per-contig embeddings for both models, per-bin metadata (coverage, GC, taxonomy, rRNA counts), Mash pairwise distances, per-bin/per-contig chimera-detection scores, ensemble operating points, and held-out / 5-fold-CV / bootstrap outputs; and (2) ZymoBIOMICS D6331 real Oxford Nanopore mock (12 MAGs, SRA SRR17913200) — mmlong2 contig/bin tables, contig-to-species gold standard, per-bin chimera labels, DNABERT-S embeddings, and a five-strain E. coli bin (bin.1.47) strain-level case study. A README.md maps every file to the corresponding figure and reported number. Analysis code: https://github.com/sunsungkim04-sys/evo2-mag. Raw CAMI2 reads: https://data.cami-challenge.org/participate ; raw D6331 reads: BioProject PRJNA804004.



