DeepARG: a deep learning approach for predicting antibiotic resistance genes from metagenomic data
收藏资源简介:
Growing concerns about increasing rates of antibiotic resistance call for expanded and comprehensive global monitoring. Advancing methods for monitoring of environmental media (e.g., wastewater, agricultural waste, food, and water) is especially needed for identifying potential resources of novel antibiotic resistance genes (ARGs), hot spots for gene exchange, and as pathways for the spread of ARGs and human exposure. Next-generation sequencing now enables direct access and profiling of the total metagenomic DNA pool, where ARGs are typically identified or predicted based on the “best hits” of sequence searches against existing databases. Unfortunately, this approach produces a high rate of false negatives. To address such limitations, we propose here a deep learning approach, taking into account a dissimilarity matrix created using all known categories of ARGs. Two deep learning models, DeepARG-SS and DeepARG-LS, were constructed for short read sequences and full gene length sequences, respectively. Evaluation of the deep learning models over 30 antibiotic resistance categories demonstrates that the DeepARG models can predict ARGs with both high precision (> 0.97) and recall (> 0.90). The models displayed an advantage over the typical best hit approach, yielding consistently lower false negative rates and thus higher overall recall (> 0.9). As more data become available for under-represented ARG categories, the DeepARG models’ performance can be expected to be further enhanced due to the nature of the underlying neural networks. Our newly developed ARG database, DeepARG-DB, encompasses ARGs predicted with a high degree of confidence and extensive manual inspection, greatly expanding current ARG repositories. The deep learning models developed here offer more accurate antimicrobial resistance annotation relative to current bioinformatics practice. DeepARG does not require strict cutoffs, which enables identification of a much broader diversity of ARGs. The DeepARG models and database are available as a command line version and as a Web service at http://bench.cs.vt.edu/deeparg.
随着人们对抗生素耐药性发生率持续攀升的担忧日益加剧,亟需开展更广泛且全面的全球监测工作。针对环境介质(如废水、农业废弃物、食品及饮用水)的监测方法迭代升级尤为关键,这有助于挖掘新型抗生素耐药基因(antibiotic resistance genes, ARGs)的潜在来源、基因交换热点区域,同时可追踪ARGs的传播路径及人类暴露途径。当下,下一代测序技术(next-generation sequencing)可直接获取并分析总宏基因组DNA库(metagenomic DNA pool),现有研究通常通过与现有数据库进行序列比对的“最佳匹配(best hits)”结果来识别或预测ARGs。但遗憾的是,该方法存在较高的假阴性率(false negatives)。为解决上述局限,本研究提出一种深度学习方法,该方法基于所有已知ARGs类别构建的差异矩阵(dissimilarity matrix)开展分析。本研究分别针对短读长序列与全长基因序列,构建了两款深度学习模型:DeepARG-SS与DeepARG-LS。针对30种抗生素耐药类别的模型评估结果显示,DeepARG系列模型可实现高精度(precision,>0.97)与高召回率(recall,>0.90)的ARGs预测。相较于传统的最佳匹配比对方法,该模型优势显著,假阴性率持续维持在较低水平,整体召回率因此更高(>0.9)。由于底层神经网络的特性,当针对占比偏低的ARG类别的数据量提升后,DeepARG模型的性能有望进一步优化。本研究新建的ARG数据库DeepARG-DB纳入了经高度可信预测且经过广泛人工核验的ARGs,大幅扩充了现有ARG数据库的收录范围。相较于当前的生物信息学(bioinformatics)分析流程,本研究开发的深度学习模型可实现更精准的抗菌耐药性注释(antimicrobial resistance annotation)。DeepARG无需设置严格的阈值(cutoffs),因此可识别更多样化的ARGs。DeepARG模型与数据库已推出命令行版本与网页服务,访问地址为http://bench.cs.vt.edu/deeparg。



