TransfIGN: A Structure-Based Deep Learning Method for Modeling the Interaction between HLA-A*02:01 and Antigen Peptides
收藏资源简介:
The intricate interaction between major histocompatibility complexes (MHCs) and antigen peptides with diverse amino acid sequences plays a pivotal role in immune responses and T cell activity. In recent years, deep learning (DL)-based models have emerged as promising tools for accelerating antigen peptide screening. However, most of these models solely rely on one-dimensional amino acid sequences, overlooking crucial information required for the three-dimensional (3-D) space binding process. In this study, we propose TransfIGN, a structure-based DL model that is inspired by our previously developed framework, Interaction Graph Network (IGN), and incorporates sequence information from transformers to predict the interactions between HLA-A*02:01 and antigen peptides. Our model, trained on a comprehensive data set containing 61,816 sequences with 9051 binding affinity labels and 56,848 eluted ligand labels, achieves an area under the curve (AUC) of 0.893 on the binary data set, better than state-of-the-art sequence-based models trained on larger data sets such as NetMHCpan4.1, ANN, and TransPHLA. Furthermore, when evaluated on the IEDB weekly benchmark data sets, our predictions (AUC = 0.816) are better than those of the recommended methods like the IEDB consensus (AUC = 0.795). Notably, the interaction weight matrices generated by our method highlight the strong interactions at specific positions within peptides, emphasizing the model’s ability to provide physical interpretability. This capability to unveil binding mechanisms through intricate structural features holds promise for new immunotherapeutic avenues.
主要组织相容性复合体(major histocompatibility complexes, MHCs)与氨基酸序列各异的抗原肽之间的复杂相互作用,在免疫应答及T细胞活性中发挥关键作用。近年来,基于深度学习(deep learning, DL)的模型已成为加速抗原肽筛选的极具潜力的工具。然而,此类模型大多仅依赖一维氨基酸序列,忽略了三维(three-dimensional, 3-D)空间结合过程所需的关键信息。本研究提出TransfIGN——一款基于结构的深度学习模型,其灵感源自我们此前开发的相互作用图网络(Interaction Graph Network, IGN)框架,并融入了Transformer(transformers)的序列信息,用于预测HLA-A*02:01与抗原肽之间的相互作用。我们的模型在包含61816条序列、9051个结合亲和力标签以及56848个洗脱配体标签的综合数据集上完成训练,在二分类数据集上取得了0.893的曲线下面积(area under the curve, AUC),优于在更大规模数据集上训练的当前最优序列模型,如NetMHCpan4.1、ANN以及TransPHLA。此外,在免疫表位数据库(Immune Epitope Database, IEDB)每周基准数据集上进行评估时,我们的预测结果(AUC=0.816)优于IEDB共识法等推荐方法(AUC=0.795)。值得注意的是,本方法生成的相互作用权重矩阵可凸显肽段特定位置的强相互作用,彰显了模型提供物理层面可解释性的能力。这种通过复杂结构特征揭示结合机制的能力,为新型免疫治疗途径带来了广阔前景。



