遇见数据集

A Simple Optimization Workflow to Enable Precise and Accurate Imputation of Missing Values in Proteomic Data Sets

收藏
Figshare2021-05-03 更新2026-04-28 收录
官方服务:

资源简介:

Missing values in proteomic data sets have real consequences on downstream data analysis and reproducibility. Although several imputation methods exist to handle missing values, no single imputation method is best suited for a diverse range of data sets, and no clear strategy exists for evaluating imputation methods for clinical DIA-MS data sets, especially at different levels of protein quantification. To navigate through the different imputation strategies available in the literature, we have established a strategy to assess imputation methods on clinical label-free DIA-MS data sets. We used three DIA-MS data sets with real missing values to evaluate eight imputation methods with multiple parameters at different levels of protein quantification: a dilution series data set, a small pilot data set, and a clinical proteomic data set comparing paired tumor and stroma tissue. We found that imputation methods based on local structures within the data, like local least-squares (LLS) and random forest (RF), worked well in our dilution series data set, whereas imputation methods based on global structures within the data, like BPCA, performed well in the other two data sets. We also found that imputation at the most basic protein quantification levelfragment levelimproved accuracy and the number of proteins quantified. With this analytical framework, we quickly and cost-effectively evaluated different imputation methods using two smaller complementary data sets to narrow down to the larger proteomic data set’s most accurate methods. This acquisition strategy allowed us to provide reproducible evidence of the accuracy of the imputation method, even in the absence of a ground truth. Overall, this study indicates that the most suitable imputation method relies on the overall structure of the data set and provides an example of an analytic framework that may assist in identifying the most appropriate imputation strategies for the differential analysis of proteins.

蛋白质组数据集的缺失值会对下游数据分析与结果可重复性产生切实影响。尽管已有多种缺失值填充方法可用于处理缺失值,但尚无单一填充方法能适配所有类型的数据集,且目前针对临床数据独立采集质谱(Data Independent Acquisition-Mass Spectrometry, DIA-MS)数据集的缺失值填充方法评估缺乏明确策略,尤其是在不同蛋白质定量层级下的评估。为梳理现有文献中各类缺失值填充策略,本研究构建了一套针对临床无标记(label-free)DIA-MS数据集的填充方法评估框架。本研究使用三个带有真实缺失值的DIA-MS数据集,在不同蛋白质定量层级下对八种带多参数的缺失值填充方法进行评估:这三个数据集分别为梯度稀释系列数据集、小型预实验数据集,以及一组配对肿瘤与间质组织的临床蛋白质组数据集。本研究发现,基于数据局部结构的填充方法(如局部最小二乘(local least-squares, LLS)与随机森林(random forest, RF))在梯度稀释系列数据集上表现优异;而基于数据全局结构的填充方法(如贝叶斯主成分分析(Bayesian Principal Component Analysis, BPCA))在其余两个数据集上表现更佳。本研究还发现,在最基础的蛋白质定量层级——即肽段片段层级——进行缺失值填充,可提升定量结果的准确性与可定量蛋白质的数量。借助该分析框架,本研究通过两个互补的小型数据集,即可快速且经济高效地评估不同缺失值填充方法,从而筛选出适配大型蛋白质组数据集的最优填充方法。即便不存在金标准(ground truth),该实验策略也能为填充方法的准确性提供可重复的验证依据。综上,本研究表明,最适配的缺失值填充方法取决于数据集的整体结构;同时本研究也提供了一套分析框架范例,可辅助筛选出适用于蛋白质差异分析的最优缺失值填充策略。

创建时间:
2021-05-03
二维码
社区交流群
二维码
科研交流群
商业服务