TernaryDB
收藏资源简介:
TernaryDB is a comprehensive dataset of 22,303 ternary complexes curated from the Protein Data Bank (PDB) to facilitate deep learning-based prediction of targeted protein degradation (TPD) complex structures. The construction involved an initial search yielding over 46,000 PDB IDs, which were then stringently filtered based on criteria such as experimental method (X-ray crystallography), resolution, R-free value, peptide chain length, and number of contacts. This process ultimately selected complexes comprising two proteins and one small molecule. The dataset features a broad chemical space, with most ligands containing fewer than 60 heavy atoms and exhibiting drug-like properties. Proteins in TernaryDB originate from 363 different species, including humans, and adequately cover PROTAC- and MG(D)-induced proteins. To prevent data leakage and ensure rigorous model assessment, the dataset was clustered by protein sequence similarity, with known PROTAC and MG(D) complexes excluded from the training set. A cluster-wise sampling strategy was implemented during training to mitigate potential biases and enhance model generalization. More Information: https://arxiv.org/abs/2502.18875



