遇见数据集

ProMeta_dataset

收藏
Zenodo2026-04-19 更新2026-05-26 收录
官方服务:

资源简介:

To evaluate the performance of the proposed ProMeta and baseline methods, we collected data from the PROTAC-DB 3.0 database. The latest release comprises 9,380 PROTAC entries, each including the SMILES representation of the compound, the UniProt identifiers of the POI and recruited E3 ligase, and degradation activity annotations including DC$_{50}$ and D$_{\max}$. We removed entries lacking SMILES, UniProt IDs, or activity labels, and validated all structures using RDKit. Binary activity labels were assigned following the protocol of Li et al.: a PROTAC is labeled as high degradation activity if DC$_{50}$ < 100 nM and D$_{\max}$ $\geq$ 80\%, and low degradation activity otherwise. To mitigate class imbalance, random majority-class down-sampling was applied independently to CRBN and VHL. The final balanced dataset comprises 860 CRBN samples and 560 VHL samples. The repository contains two CSV files: \texttt{protac\_filtered\_balanced.csv}, comprising the full balanced dataset with SMILES strings, UniProt identifiers, and binary activity labels; and \texttt{protac\_with\_seq.csv}, a subset used for model training that additionally includes the amino acid sequences of the POI and E3 ligase, with 89 entries excluded due to failure to map to a valid UniProt sequence. Molecular graph inputs were constructed from the corresponding SDF file generated via RDKit, which stores the 2D molecular structures of all PROTAC compounds in the dataset. The processed dataset and source code are available at \url{https://github.com/yufeiye-1023/ProMeta}.

提供机构:
Zenodo
创建时间:
2026-04-19
二维码
社区交流群
二维码
科研交流群
商业服务