遇见数据集

Multi-Modal CLIP-Informed Protein Editing

收藏
Zenodo2025-04-19 更新2026-05-26 收录
官方服务:

资源简介:

Proteins govern most biological functions essential for life, and achieving controllable protein editing has made great advances in probing natural systems, creating therapeutic conjugates and generating novel protein constructs. Recently, machine learning-assisted protein editing (MLPE) has shown promise in accelerating optimization cycles and reducing experimental workloads. However, current methods struggle with the vast combinatorial space of potential protein edits and cannot explicitly conduct protein editing using biotext instructions, limiting their interactivity with human feedback. To fill these gaps, we propose a novel method called ProtET for efficient CLIP-informed protein editing through multi-modality learning. Our approach comprises two stages: in the pretraining stage, contrastive learning aligns protein-biotext representations encoded by two large language models (LLMs), respectively. Subsequently, during the protein editing stage, the fused features from editing instruction texts and original protein sequences serve as the final editing condition for generating target protein sequences. This repository contains the multi-modal pretraining dataset and the model weights.

提供机构:
Zenodo
创建时间:
2025-04-19
二维码
社区交流群
二维码
科研交流群
商业服务