遇见数据集

Interpretably deep learning amyloid nucleation by massive experimental quantification of random sequences

收藏
官方服务:

资源简介:

More than 50 human diseases are characterized by the deposition of specific protein aggregates in the form of insoluble amyloid fibrils. However, only a very small number of proteins are known to form amyloids with high propensity, limiting our ability to understand, predict and engineer amyloid aggregation from sequence. Here we use a massively parallel assay to quantify the amyloid nucleation propensity of >100,000 random 20 amino acid sequences. Approximately 5% of assayed random sequences nucleate the formation of aggregates, generating a very large and diverse training dataset from which to train models to predict amyloid nucleation. We use this dataset to train CANYA, a convolution-attention hybrid neural network that predicts the propensity of any primary sequence to form amyloids. CANYA outperforms previous predictors of protein aggregation on additional random sequences and out-of-sample datasets including human disease-causing amyloids, with very stable performance across diverse prediction tasks. We adapt and extend recent advances in interpretability of genomic neural networks to elucidate CANYA’s decision-making process and learned grammar and to provide mechanistic insights into amyloid formation. Our results demonstrate the power of massive experimental random sequence-space exploration and provide an interpretable and robust neural network model for understanding, predicting and designing amyloid-forming proteins.

超过50种人类疾病的病理特征为特定蛋白质以不溶性淀粉样原纤维(amyloid fibrils)的形式沉积聚集。然而,目前已知仅极少数蛋白质具备高聚集倾向并形成淀粉样纤维,这极大限制了我们从序列层面理解、预测乃至工程化改造淀粉样聚集过程的能力。本研究采用大规模平行测定技术,对超过10万条随机20肽序列的淀粉样成核倾向开展定量表征。约5%的待测随机序列可诱导聚集体形成,由此构建出规模庞大且多样性充足的训练数据集,为淀粉样成核预测模型的训练提供了坚实基础。我们利用该数据集训练得到CANYA——一种卷积-注意力混合神经网络,可精准预测任意一级序列形成淀粉样纤维的倾向。在额外随机序列及包含人类致病淀粉样纤维的外样本数据集上,CANYA的表现均优于既往蛋白质聚集预测工具,且在各类预测任务中均展现出优异的稳定性。我们借鉴并拓展了近期基因组神经网络可解释性领域的研究进展,以此解析CANYA的决策逻辑与学习得到的序列语法,并为淀粉样纤维形成过程提供机制层面的深入见解。本研究证实了大规模实验性随机序列空间探索的研究价值,并为理解、预测与设计淀粉样形成蛋白提供了一款可解释且鲁棒的神经网络模型。

二维码
社区交流群
二维码
科研交流群
商业服务