遇见数据集

Bayesian Independent Component Analysis Recovers Pathway Signatures from Blood Metabolomics Data

收藏
Figshare2016-02-20 更新2026-04-29 收录
官方服务:

资源简介:

Interpreting the complex interplay of metabolites in heterogeneous biosamples still poses a challenging task. In this study, we propose independent component analysis (ICA) as a multivariate analysis tool for the interpretation of large-scale metabolomics data. In particular, we employ a Bayesian ICA method based on a mean-field approach, which allows us to statistically infer the number of independent components to be reconstructed. The advantage of ICA over correlation-based methods like principal component analysis (PCA) is the utilization of higher order statistical dependencies, which not only yield additional information but also allow a more meaningful representation of the data with fewer components. We performed the described ICA approach on a large-scale metabolomics data set of human serum samples, comprising a total of 1764 study probands with 218 measured metabolites. Inspecting the source matrix of statistically independent metabolite profiles using a weighted enrichment algorithm, we observe strong enrichment of specific metabolic pathways in all components. This includes signatures from amino acid metabolism, energy-related processes, carbohydrate metabolism, and lipid metabolism. Our results imply that the human blood metabolome is composed of a distinct set of overlaying, statistically independent signals. ICA furthermore produces a mixing matrix, describing the strength of each independent component for each of the study probands. Correlating these values with plasma high-density lipoprotein (HDL) levels, we establish a novel association between HDL plasma levels and the branched-chain amino acid pathway. We conclude that the Bayesian ICA methodology has the power and flexibility to replace many of the nowadays common PCA and clustering-based analyses common in the research field.

解析异质性生物样本中代谢物间的复杂相互作用,仍是一项极具挑战性的工作。本研究提出将独立成分分析(Independent Component Analysis,ICA)作为多变量分析工具,用于解读大规模代谢组学数据。具体而言,本研究采用基于平均场(mean-field)方法的贝叶斯ICA方法,可通过统计推断确定待重构的独立成分数量。相较于基于相关性的方法(如主成分分析(Principal Component Analysis,PCA)),ICA的优势在于其利用了高阶统计相关性,不仅可获取额外信息,还能以更少的成分实现更具生物学意义的数据表征。本研究将上述ICA方法应用于大规模人类血清样本代谢组学数据集,该数据集涵盖1764名研究受试者与218种已检测代谢物。通过加权富集算法对统计独立的代谢物特征源矩阵进行分析,我们发现所有成分均显著富集特定代谢通路,包括氨基酸代谢、能量代谢过程、碳水化合物代谢以及脂质代谢的特征信号。研究结果表明,人类血液代谢组由一系列独特的叠加式统计独立信号构成。此外,ICA可生成混合矩阵,用于描述每名研究受试者各独立成分的强度。将这些强度值与血浆高密度脂蛋白(High-Density Lipoprotein,HDL)水平进行关联分析后,我们发现了HDL血浆水平与支链氨基酸代谢通路间的全新关联。综上,本研究认为贝叶斯ICA方法具备足够的效力与灵活性,可替代当前该研究领域中广泛使用的主成分分析与基于聚类的诸多分析方法。

创建时间:
2016-02-20
二维码
社区交流群
二维码
科研交流群
商业服务