遇见数据集

Computational Variation: An Underinvestigated Quantitative Variability Caused by Automated Data Processing in Untargeted Metabolomics

收藏
Figshare2021-06-16 更新2026-04-28 收录
官方服务:

资源简介:

Computational tools are commonly used in untargeted metabolomics to automatically extract metabolic features from liquid chromatography-mass spectrometry (LC-MS) raw data. However, due to the incapability of software to accurately determine chromatographic peak heights/areas for features with poor chromatographic peak shape, automated data processing in untargeted metabolomics faces additional quantitative variation (i.e., computational variation) besides the well-recognized analytical and biological variations. In this work, using multiple biological samples, we investigated how experimental factors, including sample concentrations, LC separation columns, and data processing programs, contribute to computational variation. For example, we found that the peak height (PH)-based quantification is more precise when MS-DIAL was used for data processing. We further systematically compared the different patterns of computational variation between PH- and peak area (PA)-based quantitative measurements. Our results suggest that the magnitude of computational variation is highly consistent at a given concentration. Hence, we proposed a quality control (QC) sample-based correction workflow to minimize computational variation by automatically selecting PH or PA-based measurement for each intensity value. This bioinformatic solution was demonstrated in a metabolomic comparison of leukemia patients before and after chemotherapy. Our novel workflow can be effectively applied on 652 out of 915 metabolic features, and over 31% (206 out of 652) of corrected features showed distinctly changed statistical significance. Overall, this work highlights computational variation, a considerable but underinvestigated quantitative variability in omics-scale quantitative analyses. In addition, the proposed bioinformatic solution can minimize computational variation, thus providing a more confident statistical comparison among biological groups in quantitative metabolomics.

非靶向代谢组学领域通常借助计算工具,从液相色谱-质谱联用(LC-MS)原始数据中自动提取代谢特征。然而,由于软件无法精准测定峰形不佳的代谢特征的色谱峰高与峰面积,非靶向代谢组学的自动化数据处理流程,在已被广泛认知的分析变异与生物学变异之外,还会引入额外的定量变异(即计算变异)。本研究依托多份生物样本,探究了样本浓度、液相色谱分离柱、数据处理程序等实验因素对计算变异的贡献程度。例如,本研究发现,采用MS-DIAL进行数据处理时,基于峰高(peak height, PH)的定量分析精度更高。后续本研究还系统对比了基于PH与基于峰面积(peak area, PA)的定量检测中,计算变异的不同分布模式。研究结果表明,在固定浓度下,计算变异的幅度具有高度一致性。据此,本研究提出了一种基于质量控制(quality control, QC)样本的校正流程:通过自动为每个强度值匹配合适的PH或PA定量方式,以最小化计算变异。该生物信息学方案已在白血病患者化疗前后的代谢组学对比分析中得到验证。本研究提出的新型流程可有效应用于915个代谢特征中的652个,且经校正后的特征中超过31%(652个中的206个)的统计学显著性发生了显著变化。综上,本研究凸显了计算变异这一问题:它是组学规模定量分析中一类影响显著却尚未得到充分探究的定量变异来源。此外,本研究提出的生物信息学方案可有效降低计算变异,从而为定量代谢组学中生物组间的统计学比较提供更可靠的依据。

创建时间:
2021-06-16
二维码
社区交流群
二维码
科研交流群
商业服务