遇见数据集

Proteome Sampling by the HLA Class I Antigen Processing Pathway

收藏
Figshare2016-01-19 更新2026-04-29 收录
官方服务:

资源简介:

The peptide repertoire that is presented by the set of HLA class I molecules of an individual is formed by the different players of the antigen processing pathway and the stringent binding environment of the HLA class I molecules. Peptide elution studies have shown that only a subset of the human proteome is sampled by the antigen processing machinery and represented on the cell surface. In our study, we quantified the role of each factor relevant in shaping the HLA class I peptide repertoire by combining peptide elution data, in silico predictions of antigen processing and presentation, and data on gene expression and protein abundance. Our results indicate that gene expression level, protein abundance, and rate of potential binding peptides per protein have a clear impact on sampling probability. Furthermore, once a protein is available for the antigen processing machinery in sufficient amounts, C-terminal processing efficiency and binding affinity to the HLA class I molecule determine the identity of the presented peptides. Having studied the impact of each of these factors separately, we subsequently combined all factors in a logistic regression model in order to quantify their relative impact. This model demonstrated the superiority of protein abundance over gene expression level in predicting sampling probability. Being able to discriminate between sampled and non-sampled proteins to a significant degree, our approach can potentially be used to predict the sampling probability of self proteins and of pathogen-derived proteins, which is of importance for the identification of autoimmune antigens and vaccination targets.

个体的人类白细胞抗原I类(HLA class I)分子所呈递的肽谱,由抗原加工通路的各类组分以及人类白细胞抗原I类分子严苛的结合环境共同塑造。肽洗脱实验研究表明,人类蛋白质组中仅有部分子集会被抗原加工机制采样并呈递至细胞表面。本研究通过整合肽洗脱数据、抗原加工与呈递的计算机模拟预测,以及基因表达和蛋白质丰度数据,量化了各相关因子在塑造人类白细胞抗原I类肽谱过程中的作用。研究结果显示,基因表达水平、蛋白质丰度以及每个蛋白质的潜在结合肽段数量,均对采样概率具有显著影响。进一步研究发现,当蛋白质以足够量可供抗原加工机制利用时,C端加工效率以及与人类白细胞抗原I类分子的结合亲和力,将决定所呈递肽段的序列特征。在分别探究各因子的影响后,我们将所有因子整合至逻辑回归模型中,以量化它们的相对影响权重。该模型证实,在预测采样概率方面,蛋白质丰度的表现优于基因表达水平。本方法能够有效区分被采样与未被采样的蛋白质,因此可潜在用于预测自身蛋白质以及病原体衍生蛋白质的采样概率,这对于自身免疫抗原与疫苗靶点的识别具有重要意义。

创建时间:
2016-01-19
二维码
社区交流群
二维码
科研交流群
商业服务