遇见数据集

Cross-Correlation of Spectral Count Ranking to Validate Quantitative Proteome Measurements

收藏
Figshare2016-02-17 更新2026-04-29 收录
官方服务:

资源简介:

The measurement of change in biological systems through protein quantification is a central theme in modern biosciences and medicine. Label-free MS-based methods have greatly increased the ease and throughput in performing this task. Spectral counting is one such method that uses detected MS2 peptide fragmentation ions as a measure of the protein amount. The method is straightforward to use and has gained widespread interest. Additionally reports on new statistical methods for analyzing spectral count data appear at regular intervals, but a systematic evaluation of these is rarely seen. In this work, we studied how similar the results are from different spectral count data analysis methods, given the same biological input data. For this, we chose the algorithms Beta Binomial, PLGEM, QSpec, and PepC to analyze three biological data sets of varying complexity. For analyzing the capability of the methods to detect differences in protein abundance, we also performed controlled experiments by spiking a mixture of 48 human proteins in varying concentrations into a yeast protein digest to mimic biological fold changes. In general, the agreement of the analysis methods was not particularly good on the proteome-wide scale, as considerable differences were found between the different algorithms. However, we observed good agreements between the methods for the top abundance changed proteins, indicating that for a smaller fraction of the proteome changes are measurable, and the methods may be used as valuable tools in the discovery-validation pipeline when applying a cross-validation approach as described here. Performance ranking of the algorithms using samples of known composition showed PLGEM to be superior, followed by Beta Binomial, PepC, and QSpec. Similarly, the normalized versions of the same method, when available, generally outperformed the standard ones. Statistical detection of protein abundance differences was strongly influenced by the number of spectra acquired for the protein and, correspondingly, its molecular mass.

通过蛋白质定量表征生物系统的变化,是现代生物科学与医学领域的核心研究主题之一。基于质谱(Mass Spectrometry, MS)的无标记定量方法,极大提升了该研究任务的便捷性与样本通量。光谱计数(Spectral counting)便是这类方法中的一种,其以检测到的二级质谱(MS2)肽段碎片离子作为蛋白质含量的衡量指标。该方法操作简便,已获得领域内的广泛关注。此外,针对光谱计数数据分析的新型统计方法研究层出不穷,但针对这些方法的系统性评估却鲜有报道。本研究以统一的生物学输入数据为基础,探究不同光谱计数数据分析方法所得结果的一致性水平。为此,我们选取了β二项分布算法(Beta Binomial)、PLGEM、QSpec及PepC四种分析算法,对三组复杂度各异的生物学数据集展开分析。为评估各方法检测蛋白质丰度差异的能力,我们还设置了对照实验:将48种不同浓度的人源蛋白质混合物掺入酵母蛋白消化物中,以此模拟生物学意义上的蛋白丰度倍数变化。总体而言,在全蛋白质组层面,各分析方法的一致性并不理想,不同算法间存在显著差异。不过,我们观察到,在丰度变化显著的蛋白质组中,各方法间的一致性较好。这表明尽管仅在蛋白质组的较小子集内可检测到丰度变化,但如本文所述采用交叉验证策略时,这些方法可作为蛋白质组学发现-验证流程中的有效工具。基于已知组成样本的算法性能排名结果显示,PLGEM表现最优,其次依次为β二项分布算法(Beta Binomial)、PepC与QSpec。同理,若对应算法存在标准化版本,其整体性能通常优于原始标准版本。蛋白质丰度差异的统计检出能力,显著受限于为该蛋白质所采集的光谱数量,以及其对应的分子质量。

创建时间:
2016-02-17
二维码
社区交流群
二维码
科研交流群
商业服务