遇见数据集

Prediction of Peptide Reactivity with Human IVIg through a Knowledge-Based Approach

收藏
Figshare2016-01-18 更新2026-04-29 收录
官方服务:

资源简介:

The prediction of antibody-protein (antigen) interactions is very difficult due to the huge variability that characterizes the structure of the antibodies. The region of the antigen bound to the antibodies is called epitope. Experimental data indicate that many antibodies react with a panel of distinct epitopes (positive reaction). The Challenge 1 of DREAM5 aims at understanding whether there exists rules for predicting the reactivity of a peptide/epitope, i.e., its capability to bind to human antibodies. DREAM 5 provided a training set of peptides with experimentally identified high and low reactivities to human antibodies. On the basis of this training set, the participants to the challenge were asked to develop a predictive model of reactivity. A test set was then provided to evaluate the performance of the model implemented so far. We developed a logistic regression model to predict the peptide reactivity, by facing the challenge as a machine learning problem. The initial features have been generated on the basis of the available knowledge and the information reported in the dataset. Our predictive model had the second best performance of the challenge. We also developed a method, based on a clustering approach, able to “in-silico” generate a list of positive and negative new peptide sequences, as requested by the DREAM5 “bonus round” additional challenge. The paper describes the developed model and its results in terms of reactivity prediction, and highlights some open issues concerning the propensity of a peptide to react with human antibodies.

由于抗体结构具有极高的变异性,抗体-抗原(antigen)相互作用的预测极具挑战性。抗原与抗体结合的区域被称为表位(epitope)。实验数据表明,诸多抗体可与一组不同的表位产生反应(即阳性反应)。DREAM5挑战赛的第一赛道旨在探究是否存在可预测肽(peptide)/表位反应性的规则,亦即其与人类抗体结合的能力。DREAM5提供了一组经实验验证、对人类抗体具有高反应性与低反应性的肽数据集作为训练集。参赛团队需基于该训练集构建反应性预测模型。随后主办方将提供测试集,以评估已构建模型的性能表现。我们将本次挑战赛视为机器学习任务,构建了用于预测肽反应性的逻辑回归(logistic regression)模型。初始特征基于现有研究知识与数据集内的相关信息生成。我们的预测模型在本次挑战赛中取得了第二优异的成绩。我们还基于聚类(clustering)方法构建了一种可通过硅基(in-silico)模拟生成阳性与阴性新肽序列列表的方法,以响应DREAM5的‘附加赛’额外挑战赛要求。本文详述了所构建的反应性预测模型及其实验结果,并针对肽与人类抗体结合的倾向性问题,指出了若干待解决的开放性议题。

创建时间:
2016-01-18
二维码
社区交流群
二维码
科研交流群
商业服务