遇见数据集

SynGlue: Fragment-Centric Generative AI Framework for Rational Design of Protein Degraders (Dataset and Code)

收藏
Zenodo2026-03-28 更新2026-05-26 收录
官方服务:

资源简介:

This dataset accompanies the manuscript: “SynGlue: Fragment-centric modeling of modular determinants for the rational design of clinically relevant protein degraders.” The repository contains all data, models, and computational workflows required to reproduce the SynGlue framework, including large-scale curated interaction data, fragment-level representations, predictive models, and generative design outputs. --- ## 🔹 Dataset Organization and Scale The dataset is organized into modular directories: (MagnetDB/, Fragments/, Models/, Benchmark/, Generated/, Scripts/) Each component is distributed to ensure efficient storage and upload. Dataset scale:• 6.37 million protein–ligand interactions • ~2.4 million fragments • 1.94 million unique compounds • 20,129 protein targets • 6,935 benchmark compounds --- ## 🔹 Raw Data Sources and Database-wise Files The dataset integrates multiple large-scale public repositories and their processed derivatives: • DrugBank (v5.1.10) - Raw XML dump (drugbank_all_full_database.xml) - Extracted protein–ligand interaction tables - Standardized ligand SMILES and target mappings • BindingDB (Downloaded: 28-May-2023) - Raw SDF/TSV files - Binding affinity records (Ki, IC50, Kd) - Parsed ligand–target interaction matrices • ChEMBL (v33) - Raw database dump (SDF/SQL/CSV) - Bioactivity tables (activities.csv) - Target dictionary and assay metadata • STITCH (v5.0) - Raw interaction network files (protein_chemical.links) - Confidence-scored protein–chemical associations - Filtered experimentally supported interactions • BioSNAP (ChG-Miner, ChG-InterDecagon, ChG-TargetDecagon; August 2018 release) - Raw network edge lists - Chemical–gene interaction mappings - Preprocessed interaction subsets • Small Molecule Suite (aggregated dataset) - Raw ligand datasets - Standardized molecular structures - Integrated protein–ligand mappings All datasets are versioned and traceable to their original releases. --- ## 🔹 Processed Core Database (MagnetDB) • MagnetDB (Final curated dataset) - 6.37 million high-confidence protein–ligand interactions - Canonical SMILES (RDKit/OpenBabel standardized) - UniProt-mapped protein identifiers - Species-filtered datasets (Human, Mouse, Yeast) - Deduplicated interaction records - Associated assay metadata (where available) --- ## 🔹 Fragmentation and Indexing • RECAP Fragment Library - ~2.4 million terminal fragments - Fragment SMILES with attachment points (*) - Fragment-to-parent ligand mappings • TRIE Index - Serialized TRIE structure (pickle format) - Reverse SMILES indexed fragments - Hash mappings linking fragments → ligands → targets --- ## 🔹 Benchmark and Modeling Datasets • Benchmark Dataset - 6,935 compounds with experimentally validated targets - Ground truth annotations - Evaluation splits • PROTAC Dataset - DC50 and Dmax values - Potency class labels (quartile-based) - Component decomposition (warhead, linker, E3 ligand) --- ## 🔹 Generated Outputs and Model Predictions • Generated PROTAC Library - BRD4–VHL candidates - GSPT1–CRBN candidates - SMILES and structural annotations • Model Predictions - DC50 regression outputs - Dmax regression outputs - Multiclass potency classification results The predictive models are trained on publicly available degradation datasets and may reflect variability across experimental assay conditions. --- ## 🔹 Models and Computational Pipelines • Model Artifacts - Trained classification models (XGBoost, Random Forest, etc.) - Transformer-based regression models - Feature selection outputs (Boruta) • Scripts and Workflows - Data curation pipelines - SMILES standardization workflows - RECAP fragmentation scripts - TRIE indexing implementation - Model training and evaluation pipelines --- ## 🔹 Reproducibility and Usage This repository enables: • Reproduction of fragment-based target mapping • Reconstruction of the MagnetDB interaction space • Training and evaluation of SynGlue predictive models • Replication of generative PROTAC design workflows All processing steps are reproducible from raw data ingestion through fragmentation, indexing, model training, and final prediction outputs. All results reported in the associated manuscript can be reproduced using the provided datasets and scripts without requiring additional proprietary resources. --- ## 🔹 Data Provenance and Licensing All datasets are traceable to their original sources (DrugBank v5.1.10, BindingDB May 2023, ChEMBL v33, STITCH v5.0, BioSNAP 2018). Original datasets are redistributed in processed form in compliance with their respective licenses. --- This dataset supports scalable fragment-centric modeling and enables the rational design of multi-target therapeutics and protein degraders.

提供机构:
Zenodo
创建时间:
2026-03-28
二维码
社区交流群
二维码
科研交流群
商业服务