PocketVina Enables Scalable and Highly Accurate Physically Valid Docking through Multi-Pocket Conditioning
收藏资源简介:
Paper Abstract:Sampling physically valid ligand-binding poses remains a major challenge in molecular docking, particularly for unseen or structurally diverse targets. We introduce PocketVina, a fast and memory-efficient, search-based docking framework that combines pocket prediction with systematic multi-pocket exploration. We evaluate PocketVina across four established benchmarks—PDBbind2020 (timesplit and unseen), DockGen, Astex, and PoseBusters—and observe consistently strong performance in sampling physically valid docking poses. PocketVina achieves state-of-the-art performance when jointly considering ligand r.m.s.d. and physical validity (PB-valid), while remaining competitive with deep learning–based approaches in terms of r.m.s.d. alone, particularly on structurally diverse and previously unseen targets. PocketVina also maintains state-of-the-art physically valid docking accuracy across ligands with varying degrees of flexibility. We further introduce TargetDock-AI, a benchmarking dataset we curated, consisting of over 500,000 protein–ligand pairs, and a partition of the dataset labeled with PubChem activity annotations. On this large-scale dataset, PocketVina successfully discriminates active from inactive targets, outperforming a deep learning baseline while requiring significantly less GPU memory and runtime. PocketVina offers a robust and scalable docking strategy that requires no task-specific training and runs efficiently on standard GPUs, making it well-suited for high-throughput virtual screening and structure-based drug discovery.
论文摘要:在分子对接领域,采样物理合理的配体结合构象始终是一项核心挑战,尤其针对未见靶点或结构多样性较强的靶点而言。我们提出了PocketVina——一款兼具快速性与内存高效性的基于搜索的对接框架,它将口袋预测与系统性多口袋探索相结合。我们在四项通用基准数据集——PDBbind2020(时间拆分版与未见靶点版)、DockGen、Astex以及PoseBusters——上对PocketVina进行了评估,结果显示其在采样物理合理的对接构象方面始终表现优异。当同时考量配体均方根偏差(root mean square deviation, RMSD)与物理合理性(PB-valid)时,PocketVina达到了当前最优性能;而仅从RMSD指标来看,其性能可与基于深度学习的方法相媲美,尤其在结构多样性较强以及此前未见的靶点上表现突出。此外,PocketVina在不同柔性程度的配体对接任务中,仍能保持物理合理对接准确率的当前最优水平。我们还推出了TargetDock-AI——一个我们精心构建的基准数据集,包含超过50万条蛋白质-配体对,且该数据集的子集带有PubChem活性注释标签。在该大规模数据集上,PocketVina能够有效区分活性靶点与非活性靶点,其性能优于深度学习基线模型,同时所需GPU显存与运行时长均大幅降低。PocketVina提供了一种稳健且可扩展的对接策略,无需针对特定任务进行训练,且可在标准GPU上高效运行,因此非常适用于高通量虚拟筛选与基于结构的药物研发工作。




