遇见数据集

Optimization and Evaluation Datasets for PiMine

收藏
DataCite Commons2024-01-25 更新2025-04-16 收录
官方服务:

资源简介:

The protein-protein interface comparison software PiMine was developed to provide fast comparisons against databases of known protein-protein complex structures. Its application domains range from the prediction of interfaces and potential interaction partners to the identification of potential small molecule modulators of protein-protein interactions.[1] The protein-protein evaluation datasets are a collection of five datasets that were used for the parameter optimization (<em>ParamOptSet</em>), enrichment assessment (<em>Dimer597</em> set, <em>Keskin</em> set, <em>PiMineSet</em>), and runtime analyses (<em>RunTimeSet</em>) of protein-protein interface comparison tools. The evaluation datasets contain pairs of interfaces of protein chains that either share sequential and structural similarities or are even sequentially and structurally unrelated. They enable comparative benchmark studies for tools designed to identify interface similarities. In addition, we added the results of the case studies analyzed in [1] to enable readers to follow the discussion and investigate the results individually. Data Set description: The <em>ParamOptSet</em> was designed based on a study on improving the benchmark datasets for the evaluation of protein-protein docking tools [2]. It was used to optimize and fine-tune the geometric search parameters of PiMine. The <em>Dimer597</em> [3] and <em>Keskin</em> [4] sets were developed earlier. We used them to evaluate PiMine’s performance in identifying structurally and sequentially related interface pairs as well as interface pairs with prominent similarity whose constituting chains are sequentially unrelated. The <em>PiMine</em> set [1] was constructed to assess different quality criteria for reliable interface comparison. It consists of similar pairs of protein-protein complexes of which two chains are sequentially and structurally highly related while the other two chains are unrelated and show different folds. It enables the assessment of the performance when the interfaces of apparently unrelated chains are available only. Furthermore, we could obtain reliable interface-interface alignments based on the similar chains which can be used for alignment performance assessments. Finally, the <em>RunTimeSet</em> [1] comprises protein-protein complexes from the PDB that were predicted to be biologically relevant. It enables the comparison of typical run times of comparison methods and represents also an interesting dataset to screen for interface similarities. References: [1] Graef, J.; Ehrt, C.; Reim, T.; Rarey, M. Database-driven identification of structurally similar protein-protein interfaces (submitted)<br> [2] Barradas-Bautista, D.; Almajed, A.; Oliva, R.; Kalnis, P.; Cavallo, L. Improving classification of correct and incorrect protein-protein docking models by augmenting the training set. Bioinform. Adv. 2023, 3, vbad012.<br> [3] Gao, M.; Skolnick, J. iAlign: a method for the structural comparison of protein–protein interfaces. Bioinformatics 2010, 26, 2259-2265.<br> [4] Keskin, O.; Tsai, C.-J.; Wolfson, H.; Nussinov, R. A new, structurally nonredundant, diverse data set of protein–protein interfaces and its implications. Protein Sci. 2004, 13, 1043-1055.

提供机构:
Universität Hamburg
创建时间:
2024-01-25
搜集汇总
数据集介绍
Optimization and Evaluation Datasets for PiMine 数据集图片
背景与挑战
背景概述
该数据集用于优化和评估蛋白质-蛋白质界面比较工具PiMine,包含五个子集:ParamOptSet用于参数优化,Dimer597、Keskin和PiMineSet用于评估界面相似性识别能力,RunTimeSet用于运行时分析。数据集提供结构相关及不相关的界面对,支持基准测试和案例研究,总大小约18.8 GB。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务