Flexible Protein-Protein Docking Benchmark(FD1.0)
收藏资源简介:
To effectively assess the capabilities of various methods in flexible protein-protein docking, it is essential for a protein-protein docking dataset to encompass not only the structures of the heterodimer but also that of unbound monomers. Existing datasets such as DB5.5 and AB-Benchmark, while useful, are relatively limited in scale. In contrast, the Database of Interacting Protein Structures (DIPS) contains up to 42,826 binary protein complex structures but lacks the unbound state structures of the monomers. This limitation restricts its applicability to evaluations of rigid docking models rather than flexible ones. Consequently, the impact of large-scale docking datasets on methods for flexible protein-protein docking has not been thoroughly explored. To address this gap, we introduce the Flexible Protein-Protein Docking Benchmark (FD1.0), which, to our knowledge, is currently the largest dataset dedicated to flexible protein-protein docking. By providing a large and well-characterized dataset, FD1.0 aims to foster innovation in the development of flexible docking algorithms. It allows researchers to rigorously test and refine their methods, facilitating more accurate predictions of protein interactions, which are essential for understanding biological functions and designing therapeutic interventions. In our analysis of the DIPS dataset, we identified several critical issues: (1) Multiple three-dimensional structures correspond to a single protein sequence, introducing substantial noise and affecting fair comparisons among baselines, especially for models reliant on 3D structural data. (2) The DIPS training set, primarily consisting of homo-multimers, fails to capture the diversity of interface types fully. Moreover, protein-protein docking predictions are most valuable for elucidating mechanisms of protein-protein interactions (PPIs), which predominantly involve heterodimers. Homomers, often synthesized directly rather than through docking, do not accurately represent typical PPI scenarios. (3) A significant number of docking cases in DIPS involve the interaction of one polymeric protein with another, further complicating the dataset. As a cornerstone for the flexible docking dataset, it is imperative to acquire the structures of protein monomers in their unbound state. Specifically, this can be achieved through protein structure prediction methods, such as AlphaFold2, and the aggregation of structural data from sources including electron microscopy. Additionally, acknowledging the deficiencies of the DIPS dataset, several guidelines were established in the construction process of the FD1.0 dataset: (1) Each protein monomer is associated with a unique three-dimensional structure, reducing dataset noise. (2) We ensured that the similarity score (as determined by MMSeq) between docking monomers does not exceed 0.6, thereby filtering out homodimeric pairs from the dataset. (3) Unlike DIPS, a certain proportion of cases in the which dataset actually involve docking of two protein multimer. Current methods for predicting multimeric structures, such as AlphaFold Multimer, still do not achieve satisfactory results (AlphaFold3's license prohibits its use for docking purposes). However, current methods for predicting monomeric structures have reached a high level of accuracy. Therefore, we filtered out such cases, ensuring that each docking instance involves only protein monomers, guaranteeing the quality of the dataset. By adhering to these standardized construction criteria and through the collection, cleaning, and organization of data from various sources, including the Protein Data Bank and existing datasets, we compiled 3721 entries. Following the DIPS division ratio, these entries were divided into training, validation, and test sets of 3546, 98, and 77, respectively.
为有效评估各类方法在柔性蛋白质-蛋白质对接(flexible protein-protein docking)中的性能,蛋白质-蛋白质对接数据集需同时涵盖异源二聚体(heterodimer)结构与未结合单体(unbound monomers)结构。现有数据集如DB5.5和AB-Benchmark虽具备一定实用性,但规模相对有限。与之形成对比的是,相互作用蛋白质结构数据库(Database of Interacting Protein Structures,DIPS)包含多达42826个二元蛋白质复合物结构,却缺失单体的未结合态结构,这一局限使其仅适用于刚性对接模型的评估,无法支撑柔性对接方法的测试。因此,大规模对接数据集对柔性蛋白质-蛋白质对接方法的影响尚未得到充分探索。 为填补这一研究空白,我们提出柔性蛋白质-蛋白质对接基准数据集(Flexible Protein-Protein Docking Benchmark,FD1.0)——据我们所知,这是目前规模最大的柔性蛋白质-蛋白质对接专用数据集。通过提供大规模且经过充分表征的数据集,FD1.0旨在推动柔性对接算法的研发创新,帮助研究人员严谨地测试与优化其方法,助力更精准的蛋白质相互作用预测——而后者是解析生物学功能与设计治疗干预手段的核心前提。 在对DIPS数据集的分析中,我们发现了若干关键问题:(1) 单条蛋白质序列对应多个三维结构,引入了大量噪声,影响了基线模型间的公平比较,尤其是依赖三维结构数据的模型。(2) DIPS训练集主要由同源多聚体(homo-multimers)构成,未能充分覆盖界面类型的多样性。此外,蛋白质-蛋白质相互作用(protein-protein interactions,PPI)机制解析中,蛋白质-蛋白质对接预测的价值主要体现在以异源二聚体为主的场景中,而同源多聚体(homomers)通常直接通过合成获得,无需通过对接过程,无法准确代表典型的PPI场景。(3) DIPS中大量对接案例涉及一个聚合态蛋白质(polymeric protein)与另一个聚合态蛋白质(polymeric protein)的相互作用,进一步加剧了数据集的复杂性。 作为柔性对接数据集的核心基础,获取单体蛋白质的未结合态结构至关重要。具体而言,可通过AlphaFold2等蛋白质结构预测方法实现,并整合来自电子显微镜等渠道的结构数据。同时,针对DIPS数据集的诸多缺陷,我们在FD1.0数据集的构建过程中制定了若干规范:(1) 每个蛋白质单体对应唯一的三维结构,以降低数据集噪声。(2) 我们确保对接单体间的相似性得分(由MMSeq测定)不超过0.6,从而从数据集中过滤掉同源二聚体对。(3) 与DIPS不同,数据集中原本存在一定比例的两个蛋白质多聚体对接的案例。当前多聚体结构(multimeric structures)预测方法(如AlphaFold Multimer)尚未取得令人满意的效果(且AlphaFold3的使用许可禁止其应用于对接任务),而单体结构预测方法已达到较高精度,因此我们过滤掉了此类案例,确保每个对接实例仅涉及蛋白质单体,保障数据集的质量。 通过遵循这些标准化构建准则,并收集、清洗与整合来自蛋白质数据库(Protein Data Bank)及现有数据集的各类数据,我们最终得到3721条条目。参考DIPS的划分比例,将这些条目划分为训练集、验证集与测试集,三者的数量分别为3546、98和77。



