Prediction of Protein Lysine Acylation by Integrating Primary Sequence Information with Multiple Functional Features
收藏资源简介:
Liquid chromatography–tandem mass spectrometry (LC–MS/MS)-based proteomic methods have been widely used to identify lysine acylation proteins. However, these experimental approaches often fail to detect proteins that are in low abundance or absent in specific biological samples. To circumvent these problems, we developed a computational method to predict lysine acylation, including acetylation, malonylation, succinylation, and glutarylation. The prediction algorithm integrated flanking primary sequence determinants and evolutionary conservation of acylated lysine as well as multiple protein functional annotation features including gene ontology, conserved domains, and protein–protein interactions. The inclusion of functional annotation features increases predictive power oversimple sequence considerations for four of the acylation species evaluated. For example, the Matthews correlation coefficient (MCC) for the prediction of malonylation increased from 0.26 to 0.73. The performance of prediction was validated against an independent data set for malonylation. Likewise, when tested with independent data sets, the algorithm displayed improved sensitivity and specificity over existing methods. Experimental validation by Western blot experiments and LC–MS/MS detection further attested to the performance of prediction. We then applied our algorithm on to the mouse proteome and reported the global-scale prediction of lysine acetylation, malonylation, succinylation, and glutarylation, which should serve as a valuable resource for future functional studies.
基于液相色谱-串联质谱(Liquid chromatography–tandem mass spectrometry, LC–MS/MS)的蛋白质组学方法已被广泛应用于赖氨酸酰化蛋白质的鉴定。然而,此类实验手段往往无法检测到丰度较低的蛋白质,或是在特定生物样本中缺失的蛋白质。为规避上述局限,我们开发了一款用于预测赖氨酸酰化的计算方法,涵盖乙酰化、丙二酰化、琥珀酰化与戊二酰化四种修饰类型。该预测算法整合了酰化赖氨酸侧翼的一级序列决定因素、酰化位点的进化保守性,同时纳入多种蛋白质功能注释特征,包括基因本体(gene ontology, GO)、保守结构域以及蛋白质-蛋白质相互作用。相较于仅考虑序列特征的分析方案,纳入功能注释特征可提升四类酰化修饰的预测性能。例如,丙二酰化预测的马修斯相关系数(Matthews correlation coefficient, MCC)从0.26提升至0.73。我们通过独立的丙二酰化数据集对该预测性能进行了验证。同样,在使用独立数据集开展测试时,本算法相较于现有方法展现出更优异的灵敏度与特异性。通过蛋白质免疫印迹(Western blot)实验以及LC–MS/MS检测进行的实验验证,进一步证实了该预测算法的性能。随后我们将该算法应用于小鼠蛋白质组,完成了赖氨酸乙酰化、丙二酰化、琥珀酰化与戊二酰化的全域规模预测,相关成果可作为未来功能研究的宝贵资源。



