Exploring the global metaplasmidome: unravelling plasmid landscapes and the spread of antibiotic resistance genes across diverse ecosystems
收藏资源简介:
Plasmid content was predicted from assembled data already publicly available or constructed from reads for this study. The assembled data supplied by Pasolli and colleagues (Pasolli et al., 2019) , metasub consortium (Danko et al., 2020) and TARA ocean (Tully et al., 2018) were used for the human microbiome, the built environment and the marine ecosystem respectively. For assembly in the current study, reads from metagenomes were selected from two main databases. For the soil ecosystem, the metagenomes were selected from the dedicated curated database “TerrestrialMetagenomeDB” (Corrêa et al., 2020). If the metagenomes were not assembled, reads were assembled by using megahit 1.2.9 with the metalarge option (Li et al., 2015) after cleaning the data with bbduk2 (qtrim=rl trimq=28 minlen=25 maq=20 ktrim=r k=25 mink=11 and a list of adapters to remove) from the bbtools suite (https://jgi.doe.gov/data-and-tools/software-tools/bbtools/). Plasmids were predicted for each assembly by using both reference-based and reference-free approaches as described in previous works (Hilpert et al., 2021; Hennequin et al., 2022) and available on the github website (https://github.com/meb-team/PlasSuite/). The databases used for the first approach included those for chromosomes (archaea and bacteria) and plasmids from RefSeq, as well as the MOB-suite tool (Robertson and Nash, 2018), SILVA (Quast et al., 2013) and phylogenetic markers hosted by chromosomes (Wu et al., 2013). The database created for this purpose is available at this address https://github.com/meb-team/PlasSuite/?tab=readme-ov-file#1-prepare-or-download-your-databases. Two reference-free methods were applied to contigs that were not affiliated with chromosomes (discarded) or plasmids (retained in the first step): PlasFlow (Krawczyk et al., 2018) and PlasClass (Pellow et al., 2020). Previously undetected viruses were removed by using ViralVerify (https://github.com/ablab/viralVerify)(Antipov et al., 2020) that provides in parallel plasmid/non-plasmid classification. This step would also remove potential plasmid-phage elements as described by Pfeifer et al. (Pfeifer et al., 2021), but would minimise false positives. Eukaryotic contamination was removed by aligning the sequences against the NT database and human chromosomes (GRCh38) using minimap2 (Li, 2018) with -x asm5 option. Contigs mapping with 95% identity for at least 80% coverage were removed. The predicted plasmids, hereafter referred as plasmid-like sequences (PLSs), were grouped by "scientific names" (i.e. 27) such as defined in the SRA metadata (air, lake, wetland…) and subsequently named ecosystems. These ecosystems were grouped in 9 biomes (Tab Supplementary 4). The data were then dereplicated by ecosystems using cd-hit-est with a threshold of 99%. The dereplicated PLSs were then clustered using MMseqs2 (Steinegger and Söding, 2017) with 80% of coverage an 90% of identity (--min-seq-id 0.90 -c 0.8 --cov-mode 1 --cluster-mode 2 --alignment-mode 3 --kmer-per-seq-scale 0.2) to define plasmid-like clusters (PLCs). The PLC sequences are included in the file "predicted_PLC.fasta" and the main features are dercribed in the file "metadata_PLC.tsv" fasta_id: fasta identification of the PLC ecosystem: ecosystem from which the PLC originates biome: biome of the ecosystem latitude, longitude: GPS coordinate of the ecosystem length: PLC length map_markers: plasmid marker genes detected by PlasSuite (Hilpert et al., 2021) map_ncbi: PLCs present in the RefSeq plasmid database(Hilpert et al., 2021) nb_genes: Number of genes detected by Prokka implemented in PlasSuite nb_args: ARGs detected by PlasSuite plascad: results from plascad (Che et al., 2021) Antipov, D., Raiko, M., Lapidus, A., and Pevzner, P.A. (2020) MetaviralSPAdes: assembly of viruses from metagenomic data. Bioinformatics 36: 4126–4129. Che, Y., Yang, Y., Xu, X., Břinda, K., Polz, M.F., Hanage, W.P., and Zhang, T. (2021) Conjugative plasmids interact with insertion sequences to shape the horizontal transfer of antimicrobial resistance genes. Proceedings of the National Academy of Sciences 118: e2008731118. Corrêa, F.B., Saraiva, J.P., Stadler, P.F., and da Rocha, U.N. (2020) TerrestrialMetagenomeDB: a public repository of curated and standardized metadata for terrestrial metagenomes. Nucleic Acids Res 48: D626–D632. Danko, D., Bezdan, D., Afshinnekoo, E., Ahsanuddin, S., Bhattacharya, C., Butler, D.J., et al. (2020) Global Genetic Cartography of Urban Metagenomes and Anti-Microbial Resistance. bioRxiv 724526. Hennequin, C., Forestier, C., Traore, O., Debroas, D., and Bricheux, G. (2022) Plasmidome analysis of a hospital effluent biofilm: Status of antibiotic resistance. Plasmid 122: 102638. Hilpert, C., Bricheux, G., and Debroas, D. (2021) Reconstruction of plasmids by shotgun sequencing from environmental DNA: which bioinformatic workflow? Briefings in Bioinformatics 22: bbaa059. Krawczyk, P.S., Lipinski, L., and Dziembowski, A. (2018) PlasFlow: predicting plasmid sequences in metagenomic data using genome signatures. Nucleic Acids Res 46: e35. Li, D., Liu, C.-M., Luo, R., Sadakane, K., and Lam, T.-W. (2015) MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph. Bioinformatics 31: 1674–1676. Li, H. (2018) Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34: 3094–3100. Pasolli, E., Asnicar, F., Manara, S., Zolfo, M., Karcher, N., Armanini, F., et al. (2019) Extensive Unexplored Human Microbiome Diversity Revealed by Over 150,000 Genomes from Metagenomes Spanning Age, Geography, and Lifestyle. Cell 176: 649-662.e20. Pellow, D., Mizrahi, I., and Shamir, R. (2020) PlasClass improves plasmid sequence classification. PLOS Computational Biology 16: e1007781. Pfeifer, E., Moura de Sousa, J.A., Touchon, M., and Rocha, E.P.C. (2021) Bacteria have numerous distinctive groups of phage–plasmids with conserved phage and variable plasmid gene repertoires. Nucleic Acids Res 49: 2655–2673. Quast, C., Pruesse, E., Yilmaz, P., Gerken, J., Schweer, T., Yarza, P., et al. (2013) The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res 41: D590–D596. Robertson, J. and Nash, J.H.E. (2018) MOB-suite: software tools for clustering, reconstruction and typing of plasmids from draft assemblies. Microbial Genomics 4:. Steinegger, M. and Söding, J. (2017) MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature Biotechnology. Tully, B.J., Graham, E.D., and Heidelberg, J.F. (2018) The reconstruction of 2,631 draft metagenome-assembled genomes from the global oceans. Scientific Data 5: 170203. Wu, D., Jospin, G., and Eisen, J.A. (2013) Systematic Identification of Gene Families for Use as “Markers” for Phylogenetic and Phylogeny-Driven Ecological Studies of Bacteria and Archaea and Their Major Subgroups. PLoS One 8:.
本研究针对质粒含量的预测,依托已公开的组装数据,或基于本研究产生的测序读段(reads)从头组装得到的序列完成。本研究分别采用Pasolli及其团队(Pasolli等,2019)、MetaSub联盟(Danko等,2020)以及TARA海洋项目(Tully等,2018)提供的组装数据,分别对应人类微生物组、人工构建环境与海洋生态系统的分析需求。针对本研究中的组装工作,我们从两大主流数据库中选取宏基因组测序读段;其中土壤生态系统的宏基因组读段取自经过手工整理的专属数据库“TerrestrialMetagenomeDB”(Corrêa等,2020)。 若宏基因组尚未完成组装,则先利用bbtools工具集(https://jgi.doe.gov/data-and-tools/software-tools/bbtools/)中的bbduk2工具对读段进行质控:设置参数为qtrim=rl、trimq=28、minlen=25、maq=20、ktrim=r、k=25、mink=11,并移除接头序列;随后采用megahit 1.2.9并开启metalarge模式(Li等,2015)完成读段组装。 针对每一组组装序列,本研究采用基于参考序列与无参考序列两种预测方法,相关流程参照已发表文献(Hilpert等,2021;Hennequin等,2022),所用工具可从GitHub平台(https://github.com/meb-team/PlasSuite/)获取。第一种基于参考序列的预测方法所用的数据库包括RefSeq数据库中的古菌、细菌染色体及质粒序列,同时涵盖MOB-suite工具集(Robertson和Nash,2018)、SILVA数据库(Quast等,2013)以及位于染色体上的系统发育标记基因序列(Wu等,2013)。本研究专用的配套数据库可通过以下地址获取:https://github.com/meb-team/PlasSuite/?tab=readme-ov-file#1-prepare-or-download-your-databases。针对第一步中未被归类为染色体(予以丢弃)或质粒(予以保留)的重叠群(contigs),本研究采用两种无参考序列预测方法:PlasFlow(Krawczyk等,2018)与PlasClass(Pellow等,2020)。利用ViralVerify工具(https://github.com/ablab/viralVerify,Antipov等,2020)移除此前未被检出的病毒序列——该工具可同时完成质粒/非质粒序列的分类。此步骤可同时移除Pfeifer等(2021)所述的潜在质粒-噬菌体元件,从而最大限度降低假阳性结果。通过minimap2工具(Li,2018)并设置-x asm5参数,将序列与NT数据库及人类染色体GRCh38进行比对,以移除真核生物污染序列;对于比对一致性≥95%且覆盖度≥80%的重叠群,予以丢弃。本次预测得到的质粒序列,后续统称为类质粒序列(plasmid-like sequences, PLSs)。我们依据SRA元数据中定义的“科学名称”(共27类,如空气、湖泊、湿地等)对其进行分组,并将每组命名为对应的生态系统;随后将这些生态系统归为9个生物群系(详见补充表4)。随后针对每个生态系统的类质粒序列,采用cd-hit-est工具以99%的相似度阈值完成去冗余操作。去冗余后的类质粒序列进一步通过MMseqs2工具(Steinegger和Söding,2017)进行聚类,设置参数为--min-seq-id 0.90、-c 0.8、--cov-mode 1、--cluster-mode 2、--alignment-mode 3、--kmer-per-seq-scale 0.2,以覆盖度≥80%、一致性≥90%为标准,最终得到类质粒簇(plasmid-like clusters, PLCs)。 类质粒簇的序列信息存储于文件"predicted_PLC.fasta"中,其主要特征信息则记录于"metadata_PLC.tsv"文件内。 ### 元数据字段说明 1. `fasta_id`:类质粒簇的FASTA序列标识 2. `ecosystem`:类质粒簇所属的生态系统 3. `biome`:该生态系统对应的生物群系 4. `latitude, longitude`:该生态系统的GPS坐标 5. `length`:类质粒簇的序列长度 6. `map_markers`:通过PlasSuite工具检测到的质粒标记基因(Hilpert等,2021) 7. `map_ncbi`:RefSeq质粒数据库中存在的类质粒簇序列(Hilpert等,2021) 8. `nb_genes`:通过PlasSuite集成的Prokka工具检测到的基因数量 9. `nb_args`:通过PlasSuite工具检测到的抗生素抗性基因(antimicrobial resistance genes, ARGs) 10. `plascad`:plascad工具的分析结果(Che等,2021) ### 参考文献 1. Antipov D, Raiko M, Lapidus A, Pevzner PA. 2020. MetaviralSPAdes: assembly of viruses from metagenomic data. *Bioinformatics* 36: 4126–4129. 2. Che Y, Yang Y, Xu X, Břinda K, Polz MF, Hanage WP, Zhang T. 2021. Conjugative plasmids interact with insertion sequences to shape the horizontal transfer of antimicrobial resistance genes. *Proceedings of the National Academy of Sciences* 118: e2008731118. 3. Corrêa FB, Saraiva JP, Stadler PF, da Rocha UN. 2020. TerrestrialMetagenomeDB: a public repository of curated and standardized metadata for terrestrial metagenomes. *Nucleic Acids Research* 48: D626–D632. 4. Danko D, Bezdan D, Afshinnekoo E, Ahsanuddin S, Bhattacharya C, Butler DJ, et al. 2020. Global Genetic Cartography of Urban Metagenomes and Anti-Microbial Resistance. bioRxiv 724526. 5. Hennequin C, Forestier C, Traore O, Debroas D, Bricheux G. 2022. Plasmidome analysis of a hospital effluent biofilm: Status of antibiotic resistance. *Plasmid* 122: 102638. 6. Hilpert C, Bricheux G, Debroas D. 2021. Reconstruction of plasmids by shotgun sequencing from environmental DNA: which bioinformatic workflow? *Briefings in Bioinformatics* 22: bbaa059. 7. Krawczyk PS, Lipinski L, Dziembowski A. 2018. PlasFlow: predicting plasmid sequences in metagenomic data using genome signatures. *Nucleic Acids Research* 46: e35. 8. Li D, Liu CM, Luo R, Sadakane K, Lam TW. 2015. MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph. *Bioinformatics* 31: 1674–1676. 9. Li H. 2018. Minimap2: pairwise alignment for nucleotide sequences. *Bioinformatics* 34: 3094–3100. 10. Pasolli E, Asnicar F, Manara S, Zolfo M, Karcher N, Armanini F, et al. 2019. Extensive Unexplored Human Microbiome Diversity Revealed by Over 150,000 Genomes from Metagenomes Spanning Age, Geography, and Lifestyle. *Cell* 176: 649-662.e20. 11. Pellow D, Mizrahi I, Shamir R. 2020. PlasClass improves plasmid sequence classification. *PLOS Computational Biology* 16: e1007781. 12. Pfeifer E, Moura de Sousa JA, Touchon M, Rocha EPC. 2021. Bacteria have numerous distinctive groups of phage–plasmids with conserved phage and variable plasmid gene repertoires. *Nucleic Acids Research* 49: 2655–2673. 13. Quast C, Pruesse E, Yilmaz P, Gerken J, Schweer T, Yarza P, et al. 2013. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. *Nucleic Acids Research* 41: D590–D596. 14. Robertson J, Nash JHE. 2018. MOB-suite: software tools for clustering, reconstruction and typing of plasmids from draft assemblies. *Microbial Genomics* 4. 15. Steinegger M, Söding J. 2017. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. *Nature Biotechnology*. 16. Tully BJ, Graham ED, Heidelberg JF. 2018. The reconstruction of 2,631 draft metagenome-assembled genomes from the global oceans. *Scientific Data* 5: 170203. 17. Wu D, Jospin G, Eisen JA. 2013. Systematic Identification of Gene Families for Use as “Markers” for Phylogenetic and Phylogeny-Driven Ecological Studies of Bacteria and Archaea and Their Major Subgroups. *PLOS ONE* 8.



