Refined Coordinate Structures and Pocket Analyses for Lysosomal Enzymes and Monogenic Metabolic Targets
收藏资源简介:
v3.0.1 Correction Changelog: Restored GAA_splice_manifest-09-26-2026.json so that every file referenced by the MANIFEST is present in this record. No scientific content was changed. Dataset Title: Refined Coordinate Structures and Pocket Analyses for Lysosomal Enzymes and Monogenic Metabolic Targets (v3.0) Overview: Fully quality-controlled 3D structural models, geometric pocket analyses, DiffDock docking enrichment, and NIH NCBI cross-validated target records for 11 therapeutic targets in this disease area: ASPA, CBS, CLN2, G6PD, GAA, GALC, GBA1, GLA, HEXA, HEXB, IDUA. This dataset is designed for direct use in structure-based drug design campaigns by academic and industrial research groups. What's New in v3.0 (Quality Changelog): (1) 100% of structures now pass the strict steric-clash quality gate (zero non-bonded heavy-atom pairs < 1.10 Å, verified by independent reload with missing-atom completion); (2) 0 large/complex target(s) in this collection were repaired via convergent KD-tree localized clash-cluster repair and staged OpenMM minimization, and their pocket/docking analyses recomputed on the clean geometry; (3) an NCBI Entrez biological validation layer was added for every target; (4) a machine-readable manifest (CSV+JSON) with per-file MD5 checksums and per-target QC metrics was added; (5) standard file naming and engine-version provenance throughout. Included Targets: TargetNCBI Gene IDIndicationStructureClashes (<1.1 Å)Min dist (Å)Validated pocketsDocked compoundsTop reference compoundNCBI statusASPA443Canavan diseaserefined0>1.10 (no pairs)02CaffeinevalidatedCBS875Homocystinuriarefined0>1.10 (no pairs)12CaffeinevalidatedCLN21200Neuronal ceroid lipofuscinosis (Batten disease)refined0>1.10 (no pairs)32CaffeinevalidatedG6PD2539G6PD deficiency (hemolytic anemia)refined0>1.10 (no pairs)32CaffeinevalidatedGAA2548Pompe diseaserefined0>1.10 (no pairs)42CaffeinevalidatedGALC2581Krabbe diseaserefined0>1.10 (no pairs)32CaffeinevalidatedGBA12629Gaucher disease and Parkinson's riskrefined0>1.10 (no pairs)42SucrosevalidatedGLA2717Fabry diseaserefined0>1.10 (no pairs)22CaffeinevalidatedHEXA3073Tay-Sachs diseaserefined0>1.10 (no pairs)62CaffeinevalidatedHEXB3074Sandhoff diseaserefined0>1.10 (no pairs)32CaffeinevalidatedIDUA3425Mucopolysaccharidosis Irefined0>1.10 (no pairs)42Caffeinevalidated Computational Methodology: (1) Structure prediction with the Boltz-2 deep-learning co-folding engine (NVIDIA NIM API); (2) physics refinement with OpenMM (Amber14-all force field, OBC2 implicit solvent, C-alpha harmonic guide restraints k_guide = 0.5, GPU/OpenCL); (3) for outlier targets: ultra-stable staged recovery (convergent KD-tree surgical pre-cleaning + cold→stabilizing→gold minimization) and localized Jacobi-style clash-cluster repair with rigid bonded-hydrogen movement, disulfide-aware (CYS→CYX) processing, and reload verification; (4) LIGSITE-style protein-solvent-protein pocket detection (0.9 Å grid, PSP ≥ 3) with geometric druggability gates (volume 120–2500 ų, ≥ 10 lining residues); (5) DiffDock (NVIDIA NIM API) docking enrichment with reference active compounds and decoy probes — higher (less negative) confidence scores indicate more reliable poses; (6) NCBI Entrez biological validation cross-referencing every target against curated NIH records. Quality Control: Clash gate definition: zero non-bonded heavy-atom pairs closer than 1.10 Å (bonded and sequence-adjacent residues excluded), measured after independent structure reload with missing-atom completion. All 11 targets in this collection pass. Aggregate: 33 validated pockets, 22 compound-target docking results, 11/11 targets with curated NCBI records. Data Contents: per target: refined coordinate structure (PDB with hydrogens), structure QC report, geometric pocket analysis (JSON), docking enrichment results (JSON), toxicology profile report (TPR, where available), consensus toxicology (JSON, where available), ClinVar variant-to-pocket map (JSON, where available), splice/chain provenance manifest (where applicable); dataset level: machine-readable MANIFEST (CSV+JSON, per-file MD5 + per-target QC), NCBI bio-validation layer (JSON), and README with methodology. Usage Notes: PDB files load directly into Schrödinger, MOE, Discovery Studio, PyMOL, ChimeraX and RDKit/OpenBabel pipelines (standard PDB format with hydrogens). Pocket residue lists in the JSON analyses use the structure's internal residue indexing (0-based); map to your numbering via the companion structure files. Docking confidence values are DiffDock confidence scores (higher is better). Important Caveat: All structures and analyses are computational predictions. They are suitable for hypothesis generation, lead discovery and campaign prioritization, and require experimental validation before clinical or therapeutic use. FAIR Compliance: F1–F4 (persistent DOI, rich machine-readable metadata, indexed in the nexus-resonance-codex community); A1–A2 (open access over HTTPS, metadata persistence via versioned records); I1–I3 (standard PDB/JSON/CSV formats, HGNC/NCBI vocabularies, qualified cross-references); R1–R1.3 (rich QC attributes, CC BY 4.0 commercial-use license, full provenance, domain standards). Integrity & Provenance: Every file in this dataset carries an MD5 checksum recorded in the machine-readable MANIFEST (CSV+JSON). Engine versions and generation dates are embedded in file names and companion QC reports. TTT-7 audit flags are organizational conventions only. License: Creative Commons Attribution 4.0 International (CC BY 4.0) — open for academic and commercial therapeutic drug design. Principal Investigator: James Paul Trageser Affiliation: Nexus Resonance Codex ORCID: 0009-0006-6678-2908 X (Twitter): @jtrag Organization: https://github.com/Nexus-Resonance-Codex



