B-PPI-DB: A Benchmarking Dataset for Bacterial Protein-Protein Interactions
收藏资源简介:
B-PPI-DB is a benchmarking dataset designed for the training and evaluation of bacterial protein-protein interaction (PPI) prediction models. The dataset is derived from the STRING database (version 12) [1] and is specifically constructed to address class imbalance, featuring a strict 1:10 positive-to-negative ratio. B-PPI-DB Construction Methodology Positive Interactions Positive bacterial PPIs were obtained from STRING v12. To ensure high reliability, interactions were selected exclusively based on experimental evidence meeting one of two strict criteria: Criterion A: A score >= 900 in the Experiments channel. Criterion B: A score >= 900 in the Experiments Transferred channel, and a score >= 900 in the Databases channel. This secondary criterion assumes that interactions conserved across taxa and annotated in curated bacterial databases represent homologous binding relationships in bacteria. Negative Interactions Negative examples were generated by randomly pairing proteins from the unique bacterial protein pool. We ensured that no sampled negative pair appeared in STRING as interacting at any confidence level. Non-Redundancy To remove redundancy and prevent homology-driven data leakage, all bacterial proteins were clustered using MMseqs2 at a 40% sequence identity threshold (--min-seq-id 0.4 -c 0.8 --cov-mode 0). Two interactions (A, B) and (C, D) were defined as redundant if protein A clustered with C and protein B clustered with D (or vice versa). Redundant interactions were removed from both positive and negative sets. For positive pairs, we retained the interaction with the highest combined STRING score for each cluster combination. For negative pairs, a single representative was retained per cluster combination.. Dataset Statistics Total Pairs: 202,829 Positive Pairs: 18,439 Negative Pairs: 184,390 Unique Proteins: 19,810 Bacterial Taxa: 2,646 Dataset Files b_ppi_db_interactions.csv (19.24 MB)Columns: protein1, protein2, tax1, tax2, protein1_cluster, protein2_cluster, labelContains 202,829 interaction pairs (18,439 positive, 184,390 negative). b_ppi_db_protein_sequences.fa (7.5 MB)FASTA file containing all 19,810 unique bacterial proteins in the dataset. [1] Szklarczyk D, Nastou K, Koutrouli M, Kirsch R, Mehryary F, Hachilif R, et al. The STRING database in 2025: protein networks with directionality of regulation. Nucleic Acids Res. 2025;53:D730–7. https://doi.org/10.1093/nar/gkae1113.



