遇见数据集

Data_Sheet_1_microBiomeGSM: the identification of taxonomic biomarkers from metagenomic data using grouping, scoring and modeling (G-S-M) approach.docx

收藏
NIAID Data Ecosystem2026-05-01 收录
官方服务:

资源简介:

Numerous biological environments have been characterized with the advent of metagenomic sequencing using next generation sequencing which lays out the relative abundance values of microbial taxa. Modeling the human microbiome using machine learning models has the potential to identify microbial biomarkers and aid in the diagnosis of a variety of diseases such as inflammatory bowel disease, diabetes, colorectal cancer, and many others. The goal of this study is to develop an effective classification model for the analysis of metagenomic datasets associated with different diseases. In this way, we aim to identify taxonomic biomarkers associated with these diseases and facilitate disease diagnosis. The microBiomeGSM tool presented in this work incorporates the pre-existing taxonomy information into a machine learning approach and challenges to solve the classification problem in metagenomics disease-associated datasets. Based on the G-S-M (Grouping-Scoring-Modeling) approach, species level information is used as features and classified by relating their taxonomic features at different levels, including genus, family, and order. Using four different disease associated metagenomics datasets, the performance of microBiomeGSM is comparatively evaluated with other feature selection methods such as Fast Correlation Based Filter (FCBF), Select K Best (SKB), Extreme Gradient Boosting (XGB), Conditional Mutual Information Maximization (CMIM), Maximum Likelihood and Minimum Redundancy (MRMR) and Information Gain (IG), also with other classifiers such as AdaBoost, Decision Tree, LogitBoost and Random Forest. microBiomeGSM achieved the highest results with an Area under the curve (AUC) value of 0.98% at the order taxonomic level for IBDMD dataset. Another significant output of microBiomeGSM is the list of taxonomic groups that are identified as important for the disease under study and the names of the species within these groups. The association between the detected species and the disease under investigation is confirmed by previous studies in the literature. The microBiomeGSM tool and other supplementary files are publicly available at: https://github.com/malikyousef/microBiomeGSM.

随着下一代测序(next generation sequencing)技术应用于宏基因组测序(metagenomic sequencing),诸多生物环境的微生物类群(microbial taxa)相对丰度值已被精准表征。利用机器学习模型对人类微生物组(human microbiome)进行建模,有望识别出微生物生物标志物(microbial biomarkers),并助力炎症性肠病(inflammatory bowel disease)、糖尿病(diabetes)、结直肠癌(colorectal cancer)等多种疾病的诊断。本研究旨在开发一款高效的分类模型,用于分析与不同疾病相关的宏基因组数据集(metagenomic datasets),以此识别关联这些疾病的分类学生物标志物(taxonomic biomarkers),并推动疾病诊断工作。本研究提出的microBiomeGSM工具将现有分类学信息(taxonomy information)融入机器学习方法,以解决宏基因组学疾病关联数据集的分类问题。该工具基于G-S-M(Grouping-Scoring-Modeling,分组-评分-建模)方法,以物种级信息作为特征,通过关联属(genus)、科(family)、目(order)等不同分类层级的分类学特征完成分类。本研究使用4个不同疾病相关的宏基因组数据集,将microBiomeGSM的性能与多种特征选择方法(如快速基于相关性过滤(Fast Correlation Based Filter, FCBF)、选择K最优(Select K Best, SKB)、极限梯度提升(Extreme Gradient Boosting, XGB)、条件互信息最大化(Conditional Mutual Information Maximization, CMIM)、最大似然与最小冗余(Maximum Likelihood and Minimum Redundancy, MRMR)以及信息增益(Information Gain, IG))及其他分类器(如自适应提升(AdaBoost)、决策树(Decision Tree)、LogitBoost、随机森林(Random Forest))进行了对比评估。在IBDMD数据集的目级分类阶元上,microBiomeGSM取得了最优性能,其受试者工作特征曲线下面积(Area under the curve, AUC)达0.98%。microBiomeGSM的另一项重要输出为与所研究疾病相关的关键分类学类群列表,以及这些类群所包含的物种名称。所检测到的物种与目标疾病之间的关联,已得到既往文献研究的验证。microBiomeGSM工具及其他补充文件可通过以下链接公开获取:https://github.com/malikyousef/microBiomeGSM。

创建时间:
2023-11-23
二维码
社区交流群
二维码
科研交流群
商业服务