遇见数据集

NegaPPI: A Curated Structural Dataset of Negative and Decoy Protein–Protein Complexes for Machine Learning

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

NegaPPI is a curated structural dataset of negative and decoy protein–protein complexes designed for machine learning and benchmarking of protein interaction models. Because experimentally validated negative protein–protein interaction structures are inherently scarce, NegaPPI comprises three complementary classes of computational negative examples: low-confidence protein pairs derived from PrePPI-AF (Class I), partner-swapped complexes generated from structurally dissimilar protein complexes (Class II), and interface-perturbed complexes in which interface residue identities and physicochemical properties are altered while the backbone geometry is retained (Class III). The current release contains 15,000 Class I and 3075 Class II complexes. Rather than being distributed as a fixed set of structures, Class III examples are generated on the fly from positive protein dimers in the GraPPI-SSL dataset [Zenodo at https://zenodo.org/records/22946815] using the accompanying code, enabling interface perturbations to be generated dynamically during model training or evaluation. Together, these complementary negative classes capture distinct forms of structural and interaction implausibility and provide a reusable resource for developing and benchmarking computational models for protein–protein interaction recognition and binding-mode assessment. Generating Class III structures. Class III interface-perturbed complexes are generated using random_mut_foldx.py, provided alongside this dataset. The script introduces random mutations at interface residues of a given protein complex and uses FoldX to produce the corresponding mutant structures. Requirements: Python 3 (standard library: argparse, os, re, shutil, subprocess, sys) NumPy Biopython (Bio.PDB) FoldX, installed and available on the system PATH or in the working directory Usage----- # run inside the `PPI` conda environment (requires numpy and biopython) python random_mut_foldx.py --pdb xxx.pdb --receptor A --ligand B # multi-chain sides, 3 mutations per side, 10 variants, custom output directory python random_mut_foldx.py --pdb complex.pdb --receptor H,L --ligand A \ --num-mut-on-each 3 --n-variants 10 --out-dir ./mutants --seed 0 Mutation strings follow the {chain}_{wt}{position}{mut} convention used throughout this project (e.g., A_Y87R,B_V166L); these are translated to FoldX's {wt}{chain}{position}{mut} format when the individual_list.txt file is written. FoldX must be installed and available on the system PATH or present in the working directory.

提供机构:
Zenodo
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务