GraPPI-SSL: A Curated Protein–Protein Complex Dataset for Self-Supervised Representation Learning
收藏资源简介:
GraPPI-SSL is a curated protein–protein complex dataset developed for self-supervised representation learning of protein interactions. The dataset integrates experimentally determined protein heterodimer complexes from the Protein Data Bank (PDB) with curated protein–protein binding-affinity data mainly from PPB-Affinity set, followed by systematic quality control, structural standardization, and redundancy filtering. In total, it includes 3099 structures with matched binding affinities. It provides residue-level structural context and well-defined interacting protein partners for training machine-learning models to learn representations of protein complexes without requiring task-specific labels. GraPPI-SSL was developed as the self-supervised pretraining dataset for GraPPI and is released as a reusable resource for protein representation learning and other structure-based machine-learning applications.



