Machine learning on large-scale proteomics data identifies tissue- and cell type-specific proteins
收藏资源简介:
Using data from 183 public human data sets from PRIDE, a machine learning model was trained to identify tissue and cell-type specific protein patterns. PRIDE projects were searched with ionbot and tissue/cell type annotation was manually added. Data from physiological samples were used to train a Random Forest model on protein abundances to classify samples into tissues and cell types. Subsequently, a one-vs-all classification and feature importance were used to analyse the most discriminating protein abundances per class. Based on protein abundance alone, the model was able to predict tissues with 98% accuracy, and cell types with 99% accuracy. The F-scores describe a clear view on tissue-specific proteins and tissue-specific protein expression patterns. In-depth feature analysis shows slight confusion between physiologically similar tissues, demonstrating the capacity of the algorithm to detect biologically relevant patterns. These results can in turn inform downstream uses, from identification of the tissue of origin of proteins in complex samples such as liquid biopsies, to studying the proteome of tissue-like samples such as organoids and cell lines
本研究依托PRIDE(PRIDE)数据库的183套公开人类数据集,训练了一款机器学习模型,用于识别组织与细胞类型特异性蛋白表达模式。研究使用ionbot(ionbot)对PRIDE数据库的数据集进行检索,并手动添加了组织/细胞类型注释信息。随后,选取生理样本的蛋白丰度数据,训练随机森林(Random Forest)分类模型,以将样本划分为对应组织与细胞类型。继而采用一对多(one-vs-all)分类法与特征重要性分析方法,对每一类样本中最具区分度的蛋白丰度特征展开解析。仅基于蛋白丰度数据,该模型即可实现98%的组织分类准确率与99%的细胞类型分类准确率。F值(F-score)可清晰呈现组织特异性蛋白及其表达模式的特征。深入的特征分析结果显示,生理特性相似的组织之间仅存在轻微分类混淆,证实了该算法能够检测到具有生物学意义的蛋白表达模式。上述研究结果可进一步为下游应用提供支撑:从液体活检等复杂样本中蛋白组织来源的鉴定,到类器官、细胞系等类组织样本的蛋白质组学研究,均具备应用价值。



