A Systematic Interrogation of AI Models for Prediction of Orthosteric and Allosteric Binding: Unmasking Allosteric Grammar with Explainable AI
收藏资源简介:
This Zenodo archive provides the datasets, model prediction outputs, evaluation results, figures, and scripts used in the study “Benchmarking AI-Based Co-Folding and Docking Models for Predicting Structures of Orthosteric and Allosteric Ligand–Protein Complexes: Decoding the Allosteric Blind Spot Using a Landscape-Guided Interpretable AI Framework.” The archive mirrors the directory structure required by the analysis pipeline and contains two datasets: an Orthosteric (Main) Dataset derived from the DynamicBind benchmark and an Allosteric Dataset derived from the Dunbrack Lab’s KinCoRe resource. All results reported in the manuscript are derived from these datasets. Protein and ligand structures are not included in this archive. All structures are publicly available from the RCSB Protein Data Bank and can be automatically downloaded using the provided scripts. This record contains the normalized per-model predicted ligand-protein complex structures for all predictors (AlphaFold3, Chai, Boltz, Protenix, DynamicBind) across the ASD and PLA datasets, together with the curated reference structures used for scoring. Contents: - plb_bench_data/ One standardized mmCIF per predicted model (model_000.cif, model_001.cif, ...), organized by predictor and dataset: af3, af3_pla, boltz, boltz_pla, chai, chai_pla, proteinx, protenix_pla, dynamicbind_pla, dynamicbind_asd, dynamicbind_asd_v2. - references_ref_cifs/ Curated reference structures used as scoring targets, including the corrected references for the DynamicBind PLA set. These are the cleaned, uniform outputs the benchmark was evaluated on. The raw, pre-normalization prediction outputs (the original files as each tool produced them, including per-tool confidence/score files, logs, and multiple ranked poses; ~200 GB total) are available from the authors upon request.



