Pilot datasets for SEQUESTRA: reference-benchmarked de novo protein design in catalytic- and non-catalytic-reference modes
收藏资源简介:
This dataset contains the complete computational outputs of two 100-design pilot applications of SEQUESTRA, a reference-benchmarked workflow for de novo small-molecule-binding protein design. The deposited projects evaluate both supported workflow modes: a catalytic-reference pilot based on dihydropteroate synthase bound to pterin-6-yl-methyl-monophosphate (PDB 2VEG; ligand CCD PMM) and a non-catalytic-reference pilot based on saxiphilin bound to saxitoxin (PDB 6O0F; ligand CCD 9SL). For each pilot, BoltzGen generated 100 protein designs of 180-200 residues. Designs were ranked solely by BoltzGen affinity_probability_binary, and the top 20% (20 designs per project) were retained using exact ceiling-rounded selection. The archive includes the initial reference inputs, complete rankings, shortlisted protein-ligand complexes, binder-only structures, selection manifests, model inputs, raw predictions, checkpoints, execution logs, consolidated decision tables, scalable SVG plots, FASTA files and checksum-based reproducibility manifests. In the catalytic-reference workflow, DLKcat and CatPred were used for reference-relative catalytic-liability screening. A candidate was excluded only when its predicted kcat was strictly higher and its predicted Km strictly lower than the corresponding reference predictions. None of the 20 shortlisted designs satisfied both exclusion conditions: 19 displayed a higher-kcat-only partial-liability signal and one satisfied neither adverse condition. All 20 therefore advanced to Boltz-2 affinity benchmarking. The reference affinity probability was 0.854538, whereas the highest candidate probability was 0.500853; consequently, no candidate exceeded the strict reference benchmark and structural-confidence assessment was skipped. In the non-catalytic-reference workflow, catalytic-liability screening was intentionally disabled. Boltz-2 benchmarking identified one of the 20 shortlisted designs, design_specification_51, with an affinity probability of 0.471851, as exceeding the reference probability of 0.456980. The candidate passed the configured overall-confidence and ligand-ipTM criteria but did not satisfy the complex-ipLDDT criterion: confidence score 0.753771 against a threshold of 0.70, ligand ipTM 0.793945 against 0.60, and complex ipLDDT 0.564893 against 0.70. Accordingly, no candidate passed the complete structural-confidence gate. For the non-catalytic pilot, the engineered terminal expression-tag sequence SNSLEVLFQ was manually removed from PDB 6O0F chain A before project creation because those residues are not part of the canonical saxiphilin sequence. The processed chain therefore ends at canonical residue Cys825. This modification affected only the engineered expression tag. No canonical biological residue was intentionally removed. The pilot calculations were completed using SEQUESTRA development releases through v0.11.3. SEQUESTRA v1.0.0 is the corresponding consolidated public software release. The deposited results are computational predictions intended for workflow validation and candidate prioritization. They do not constitute experimental evidence of binding, catalytic activity, non-catalytic behavior or biological function. SEQUESTRA source code, installation instructions, and a comprehensive beginner-oriented tutorial are available at: https://github.com/Olanrewaju-Durojaye/SEQUESTRA



