Dataset and Pre-trained Models for: Assay-Aware Data Curation and Nested Scaffold Cross-Validation for Large-Scale hERG Cardiotoxicity Prediction: An Operational Benchmarking Study
收藏资源简介:
This repository contains the curated datasets, extracted molecular features, and pre-trained machine learning models supporting the manuscript titled "Assay-Aware Data Curation and Nested Scaffold Cross-Validation for Large-Scale hERG Cardiotoxicity Prediction: An Operational Benchmarking Study". Files included: Curated Datasets: Processed hERG toxicity records with multi-source proxy flags and confidence weights. External Validation: The temporal out-of-distribution (OOD) blind test set collected from ChEMBL. Pre-trained Models: Optimized XGBoost models (.joblib) and selected feature mappings ready for deployment. For the full source code, preprocessing scripts, and exact environment configurations, please visit our official GitHub repository: https://github.com/zhii9527/herg-benchmark



