MCAM: Multiple Clustering Analysis Methodology for Deriving Hypotheses and Insights from High-Throughput Proteomic Datasets
收藏资源简介:
Advances in proteomic technologies continue to substantially accelerate capability for generating experimental data on protein levels, states, and activities in biological samples. For example, studies on receptor tyrosine kinase signaling networks can now capture the phosphorylation state of hundreds to thousands of proteins across multiple conditions. However, little is known about the function of many of these protein modifications, or the enzymes responsible for modifying them. To address this challenge, we have developed an approach that enhances the power of clustering techniques to infer functional and regulatory meaning of protein states in cell signaling networks. We have created a new computational framework for applying clustering to biological data in order to overcome the typical dependence on specific a priori assumptions and expert knowledge concerning the technical aspects of clustering. Multiple clustering analysis methodology (‘MCAM’) employs an array of diverse data transformations, distance metrics, set sizes, and clustering algorithms, in a combinatorial fashion, to create a suite of clustering sets. These sets are then evaluated based on their ability to produce biological insights through statistical enrichment of metadata relating to knowledge concerning protein functions, kinase substrates, and sequence motifs. We applied MCAM to a set of dynamic phosphorylation measurements of the ERRB network to explore the relationships between algorithmic parameters and the biological meaning that could be inferred and report on interesting biological predictions. Further, we applied MCAM to multiple phosphoproteomic datasets for the ERBB network, which allowed us to compare independent and incomplete overlapping measurements of phosphorylation sites in the network. We report specific and global differences of the ERBB network stimulated with different ligands and with changes in HER2 expression. Overall, we offer MCAM as a broadly-applicable approach for analysis of proteomic data which may help increase the current understanding of molecular networks in a variety of biological problems.
蛋白质组学技术的持续进步,极大提升了在生物样本中获取蛋白质水平、状态及活性相关实验数据的能力。例如,当前针对受体酪氨酸激酶(receptor tyrosine kinase)信号网络的研究,已可在多种实验条件下捕获数百至数千种蛋白质的磷酸化状态。然而,目前对于这些蛋白质修饰的多数功能,以及负责执行这些修饰的酶类,仍知之甚少。为应对这一挑战,我们开发了一种能够增强聚类技术(clustering techniques)效能的方法,以推断细胞信号网络中蛋白质状态的功能与调控意义。为克服传统聚类分析对特定先验假设(a priori assumptions)及聚类技术细节相关专业知识的固有依赖,我们构建了全新的计算框架,用于将聚类分析应用于生物数据。多聚类分析方法(Multiple Clustering Analysis Methodology,MCAM)以组合式策略整合多种数据变换、距离度量、集合规模参数与聚类算法,生成一系列聚类集。随后,我们基于这些聚类集通过对与蛋白质功能、激酶底物(kinase substrates)及序列基序(sequence motifs)相关知识元数据的统计富集能力,评估其能否产生有价值的生物学见解。我们将MCAM应用于ERRB网络的一组动态磷酸化检测数据,以探究算法参数与可推断的生物学意义之间的关联,并报告了颇具参考价值的生物学预测结果。此外,我们将MCAM应用于ERBB网络的多组磷酸蛋白质组数据集(phosphoproteomic datasets),借此可对比该网络中磷酸化位点的独立检测结果与不完全重叠的检测数据。我们报告了不同配体(ligands)刺激及HER2表达改变时,ERBB网络的特异性与全局性差异。总体而言,我们将MCAM作为一种普适性的蛋白质组数据分析方法推出,该方法或有助于增进人们在多种生物学问题中对分子网络的现有认知。



