Property-Dependent Scoring Bias in Virtual Screening Validation: Data and Code
收藏资源简介:
Complete dataset and analysis code for a dual scoring function virtual screening campaign against the human BCL-2 promoter G-quadruplex (G4) structure (PDB: 2F8U), with retrospective validation against the ROBIN (Repository Of BINders) experimental binding dataset. This dataset accompanies the manuscript:Property-Dependent Scoring Bias in Virtual Screening Validation: Evidence from Dual Scoring Function Cross-Validation of BCL-2 G-Quadruplex Ensemble DockingTanigawa M. Journal of Computer-Aided Molecular Design (submitted). Contents Receptor structures: 10 NMR conformer ensemble (PDB: 2F8U) prepared for AutoDock Vina (PDBQT format) AutoDock Vina docking data: Ensemble docking scores for 114 compounds across all 10 NMR models (1,140 calculations), including consensus rankings, per-model raw scores, and rank stability metrics rDock cross-validation data: Ensemble docking scores for 113 compounds across all 10 NMR models (1,130 calculations), including consensus rankings and inter-model concordance Cross-scoring comparison: Vina vs. rDock consensus rankings for 38 overlapping compounds, with molecular descriptor analysis of scoring function disagreement Compound annotations: Comprehensive metadata for 114 compounds (33 fields) and Online Resource 1 with 251 compounds (35 fields), including SMILES, molecular properties, pharmacological classes, and dual scoring function rankings Docking poses: PDBQT files for 114 compounds and 33 negative controls across 10 NMR models (AutoDock Vina) Sigma-scaling perturbation analysis: Coordinate perturbation results for both scoring functions (20 compounds × 5 sigma levels × 3 replicates = 300 calculations per scoring function) Discrimination analysis: G4-specific ligands vs. FDA-approved negative controls — Vina: AUROC = 0.748, p = 0.010 (n = 13 vs. 33); rDock: AUROC = 0.602, p = 0.162 (n = 11 vs. 32) ROBIN retrospective validation: Ensemble docking of 111 experimentally confirmed BCL-2 G4 SMM binders and 700 negative controls (500 random + 200 property-matched) with both scoring functions, demonstrating AUROC ≈ 0.5 for both Vina and rDock, establishing that the curated library discrimination reflects property-dependent scoring bias rather than target-specific recognition Supplementary figures: Online Resources 2–3 in PDF and PNG formats Analysis pipeline: 49 Python scripts (35 curated library + 10 ROBIN validation + 4 figure generation) for full reproducibility Key Findings Metric Vina rDock Compounds screened (single model) 7,469 7,296 Ensemble-docked compounds 114 113 Kendall's W (inter-model concordance) 0.651 0.454 AUROC (curated: G4-specific vs. negative controls) 0.748 0.602 AUROC (ROBIN: vs. random negatives) 0.466 0.493 AUROC (ROBIN: vs. property-matched negatives) 0.473 0.510 Cross-scoring Spearman ρ (38 compounds) 0.263 (p = 0.111) Software AutoDock Vina (52ec525-mod) rDock 2013.1 fpocket 4.2 Open Babel 3.1.0 RDKit 2023.09+ Python 3.8+ with NumPy, Pandas, SciPy, BioPython, Matplotlib, Seaborn Version History v2.0: Added ROBIN retrospective validation data (robin_validation/ directory), rDock negative control ensemble docking, figure generation scripts, and supplementary figures v1.0: Initial release with curated library docking data and dual scoring function cross-validation



