遇见数据集

Data and code for “Uncertainty-aware integration of multi-target prediction and dual-state docking prioritizes generated 5-HT2A candidates

收藏
Mendeley Data2026-09-08 收录
官方服务:

资源简介:

This dataset provides the data, code, trained models, computational outputs, and provenance documentation supporting the manuscript “Uncertainty-aware integration of multi-target prediction and dual-state docking prioritizes generated 5-HT2A candidates”. The release includes raw and processed ChEMBL-derived activity data, the strict binding-Ki dataset comprising 14,587 target-structure records across 5-HT2A, 5-HT2B, 5-HT2C, and 5-HT1A, scaffold-disjoint train/validation/test partitions, LightGBM and Random Forest model outputs, per-repeat predictions and uncertainty summaries, applicability-domain analyses, archived molecular prior and reinforcement-learning checkpoints, fixed-seed generative-model resampling results, candidate-level strict rescoring and chemical-quality assessments, dual-state docking and native-ligand redocking outputs, and source data for the manuscript figures and tables. Historical reinforcement-learning checkpoints and the archived candidate pool are preserved unchanged as provenance artifacts. The original stochastic RL training seeds and historical candidate-pool sampling seed were not recorded and are not retrospectively inferred. A separate deterministic reproducibility layer provides seed-controlled rerun implementations for future RL training and candidate sampling, with isolated output paths so that historical manuscript artifacts cannot be overwritten. The archive also contains environment specifications, a data dictionary, model and provenance documentation, a complete file manifest, and SHA-256 checksums for integrity verification. The four molecules cand_77, cand_164, cand_76, and cand_134 are computationally prioritized candidates and have not been experimentally validated.

本数据集提供支撑论文《不确定性感知的多靶点预测与双态对接融合方法优先筛选生成的5-HT2A受体候选化合物》的实验数据、代码、训练好的模型、计算输出结果及溯源文档。 本次发布包含原始及经处理的、源自ChEMBL的活性数据;包含覆盖5-HT2A、5-HT2B、5-HT2C及5-HT1A四种靶点的严格结合Ki数据集,共计14587条靶点-结构记录;包含支架不相交的训练/验证/测试集划分;包含LightGBM与随机森林(Random Forest)模型的输出结果、每轮重复的预测值与不确定性汇总、适用域分析结果;包含存档的分子先验模型与强化学习(Reinforcement Learning, RL)检查点、固定种子的生成模型重采样结果;包含候选分子层面的严格重打分与化学质量评估结果;包含双态对接与天然配体重新对接的输出结果;以及支撑论文图表与表格的源数据。 历史强化学习检查点与存档的候选分子池将作为溯源工件完整保留。原始随机强化学习训练种子与历史候选池采样种子未被记录,也无法通过回溯推断得到。本数据集提供独立的确定性可复现性模块,该模块支持种子可控的强化学习训练与候选分子采样重运行,且设有独立的输出路径,避免覆盖原论文相关的历史溯源工件。 该存档还包含环境配置文件、数据字典、模型与溯源文档、完整的文件清单,以及用于完整性校验的SHA-256校验和。其中cand_77、cand_164、cand_76与cand_134这四个分子为计算优先筛选出的候选化合物,尚未经过实验验证。

创建时间:
2026-08-26
二维码
社区交流群
二维码
科研交流群
商业服务