遇见数据集

B-PPI-DB: A Benchmarking Dataset for Bacterial Protein-Protein Interactions

收藏
Zenodo2025-12-17 更新2026-05-26 收录
官方服务:

资源简介:

B-PPI-DB is a benchmarking dataset designed for the training and evaluation of bacterial protein-protein interaction (PPI) prediction models. The dataset is derived from the STRING database (version 12) [1] and is specifically constructed to address class imbalance, featuring a strict 1:10 positive-to-negative ratio. B-PPI-DB Construction Methodology Positive Interactions Positive bacterial PPIs were obtained from STRING v12. To ensure high reliability, interactions were selected exclusively based on experimental evidence meeting one of two strict criteria: Criterion A: A score >= 900 in the Experiments channel. Criterion B: A score >= 900 in the Experiments Transferred channel, and a score >= 900 in the Databases channel. This secondary criterion assumes that interactions conserved across taxa and annotated in curated bacterial databases represent homologous binding relationships in bacteria. Negative Interactions Negative examples were generated by randomly pairing proteins from the unique bacterial protein pool. We ensured that no sampled negative pair appeared in STRING as interacting at any confidence level. Non-Redundancy To remove redundancy and prevent homology-driven data leakage, all bacterial proteins were clustered using MMseqs2 at a 40% sequence identity threshold (--min-seq-id 0.4 -c 0.8 --cov-mode 0). Two interactions (A, B) and (C, D) were defined as redundant if protein A clustered with C and protein B clustered with D (or vice versa). Redundant interactions were removed from both positive and negative sets. For positive pairs, we retained the interaction with the highest combined STRING score for each cluster combination. For negative pairs, a single representative was retained per cluster combination.. Dataset Statistics Total Pairs: 202,829 Positive Pairs: 18,439 Negative Pairs: 184,390 Unique Proteins: 19,810 Bacterial Taxa: 2,646 Dataset Files b_ppi_db_interactions.csv (19.24 MB)Columns: protein1, protein2, tax1, tax2, protein1_cluster, protein2_cluster, labelContains 202,829 interaction pairs (18,439 positive, 184,390 negative). b_ppi_db_protein_sequences.fa (7.5 MB)FASTA file containing all 19,810 unique bacterial proteins in the dataset. [1] Szklarczyk D, Nastou K, Koutrouli M, Kirsch R, Mehryary F, Hachilif R, et al. The STRING database in 2025: protein networks with directionality of regulation. Nucleic Acids Res. 2025;53:D730–7. https://doi.org/10.1093/nar/gkae1113.

B-PPI-DB是一款专为细菌蛋白质相互作用(Protein-Protein Interaction, PPI)预测模型的训练与评估打造的基准数据集。该数据集源自STRING数据库(版本12)[1],其构建核心目标为解决类别不平衡问题,采用严格的1:10正负样本比例。 ## B-PPI-DB 构建方法 ### 正样本交互对 正样本细菌PPI取自STRING v12数据库。为确保交互对的高可靠性,仅选取符合以下两项严格标准之一的实验证据支撑的相互作用: 1. 标准A:实验证据通道(Experiments channel)得分≥900。 2. 标准B:实验转移证据通道(Experiments Transferred channel)得分≥900,且数据库证据通道(Databases channel)得分≥900。该二级标准假定,跨类群保守且经人工手工注释(curated)的细菌数据库所标注的相互作用,对应细菌体内的同源结合关系。 ### 负样本交互对 负样本通过从唯一细菌蛋白质池中随机配对蛋白质生成。我们确保所有采样得到的负样本对,未在STRING数据库中以任何置信水平被标注为存在相互作用。 ### 非冗余性处理 为去除冗余交互对并避免同源性导致的数据泄露,所有细菌蛋白质使用MMseqs2工具以40%序列同一性阈值进行聚类(参数:--min-seq-id 0.4 -c 0.8 --cov-mode 0)。若蛋白质A与C聚类、蛋白质B与D聚类(或反之),则将交互对(A,B)与(C,D)定义为冗余交互对。 随后从正负样本集中移除冗余交互对:对于正样本对,保留每类聚类组合中STRING综合得分最高的交互对;对于负样本对,每类聚类组合仅保留一个代表性样本对。 ## 数据集统计 - 总交互对:202,829 - 正样本对:18,439 - 负样本对:184,390 - 唯一蛋白质数:19,810 - 细菌类群数:2,646 ## 数据集文件 1. `b_ppi_db_interactions.csv`(19.24 MB):列信息包括protein1、protein2、tax1、tax2、protein1_cluster、protein2_cluster、label,共包含202,829条交互对(其中正样本对18,439条,负样本对184,390条)。 2. `b_ppi_db_protein_sequences.fa`(7.5 MB):FASTA格式文件,包含数据集中全部19,810条唯一细菌蛋白质序列。 [1] Szklarczyk D, Nastou K, Koutrouli M, Kirsch R, Mehryary F, Hachilif R, 等. 2025年版STRING数据库:带调控方向性的蛋白质网络. 核酸研究, 2025;53:D730–7. https://doi.org/10.1093/nar/gkae1113.

提供机构:
Zenodo
创建时间:
2025-12-17
二维码
社区交流群
二维码
科研交流群
商业服务