Prebuilt Sylph Databases and Taxonomy Files from GTDB Release 220 - Dereplicated at 95% and 99% ANI for StrainSpy Analyses
收藏资源简介:
This dataset contains prebuilt Sylph databases generated from GTDB release 220 (r220) that were used in StrainSpy analyses. Two Sylph databases are provided with different dereplication thresholds: 1. GTDB Dereplicated at 95% ANI A genome collection dereplicated at 95% ANI, containing 113,104 genomes. This represents a species-level database designed to provide broad taxonomic coverage while maintaining a relatively compact database size. This is the same as the GTDB r220 database generated using the -c 200 setting for increased sensitivity in the original Sylph publication and also available here 2. GTDB Dereplicated at 99% ANI A GTDB r220 genome collection dereplicated at 99% ANI, containing 235,601 genomes. This higher-resolution database was generated to retain substantially more closely related genomic diversity than the 95% ANI database. The purpose of the 99% ANI database is to improve strain-level resolution by retaining additional within-species genomic diversity. This is particularly relevant for strain-level microbiome-wide association studies (MWAS) using StrainSpy, where it is observed that differences between closely related strains between groups may represent biologically meaningful variation. The ANI thresholds (95%,99%) refer to dereplication of the GTDB genome collection during database construction. They do not represent ANI cutoffs applied to query metagenomes and do not alter Sylph ANI estimates. Instead, they determine the amount of closely related genomic diversity retained in the database. Database construction: The GTDB derep 99 database was dereplicated using the greedy clustering approach implemented in genoreps. This was constructed using Sylph v0.9.0 using the -c 200 setting, providing increased database sensitivity. Contents: Database: gtdb-r220-c200-dbv1.syldb, Taxonomy file: sylph_DB_taxonomy_95.tsv Database: gtdb_ordered_99.syldb, Taxonomy file: sylph_DB_taxonomy.tsv The *.syldb files are Sylph databases and can be used directly with Sylph to estimate ANI between query metagenomes and GTDB genome collections. Example Sylph query: sylph query gtdb_ordered_99.syldb -u --read-seq-id 99.9 -t 100 \ -1 <sample_R1.fastq.gz> \ -2 <sample_R2.fastq.gz> \ -o <output.tsv> The accompanying sylph_DB_taxonomy_*.tsv files contain taxonomic assignments corresponding to each genome in the database and can be imported into StrainSpy for downstream strain-level association analyses. For Sylph:https://github.com/bluenote-1577/sylph For StrainSpy:https://github.com/gtonkinhill/strainspy



