ProtCLIP: Function-Informed Protein Multi-Modal Learning
收藏资源简介:
Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model.
旨在对齐蛋白质序列与生物学描述的多模态预训练范式,已习得通用蛋白质表征,并在各类下游应用中取得了亮眼表现。然而,由于未能有效利用对齐后的蛋白质-文本配对数据,且缺乏有效的功能感知预训练范式,这类工作仍无法复刻语言监督下的视觉基础模型所取得的卓越成功。为解决上述问题,本文构建了一个名为ProtAnno的大规模蛋白质-文本配对数据集,其采用属性驱动的采样策略,并提出了一种全新的功能感知蛋白质预训练范式。具体而言,该采样策略基于样本置信度与属性覆盖度确定选择概率,在面对大规模噪声数据时平衡了数据质量与数据规模。此外,受蛋白质特异性功能机制重要性的启发,所提出的范式通过两个分段级预训练目标,显式建模了蛋白质的静态与动态功能区段,以功能感知的方式注入细粒度信息。依托上述所有创新,我们开发了ProtCLIP——一种可全面表征功能感知蛋白质嵌入的多模态基础模型。在涵盖5大类共22种不同蛋白质基准任务中,包括蛋白质功能分类、突变效应预测、跨模态转换、语义相似度推理以及蛋白质-蛋白质相互作用预测,我们的ProtCLIP始终取得了SOTA(State-of-the-Art)性能,其中在5项跨模态转换基准任务中平均提升达75%,在GO-CC与GO-BP蛋白质功能预测任务中分别实现了59.9%与39.7%的显著提升。实验结果验证了ProtCLIP作为蛋白质多模态基础模型所具备的非凡潜力。



