NovoMCP Open Corpus (lite)
收藏资源简介:
122,454,458 PubChem compounds with RDKit physicochemical descriptors, ~30 ADMET/toxicity predictions (NovoMCP models, trained on the public Therapeutics Data Commons benchmark), and PAINS structural-alert flags. One row per PubChem CID, stored as Apache Parquet (Snappy) for direct in-place querying (Athena, DuckDB, PyArrow, pandas). An analysis-ready cheminformatics substrate for drug discovery: virtual screening, drug-likeness (QED) filtering, similarity search, chemical-space exploration, and training ADMET models at scale — so you don't recompute descriptors and property predictions acr...
本数据集包含122,454,458个PubChem化合物,配套RDKit理化描述符、约30项ADMET(吸收、分布、代谢、排泄与毒性,Absorption, Distribution, Metabolism, Excretion, Toxicity)相关预测结果(基于公开治疗数据公用库 (Therapeutics Data Commons) 训练的NovoMCP模型生成)以及PAINS(泛化验干扰化合物,Pan-Assay Interference Compounds)结构警示标记。每条数据行对应一个PubChem CID,以Apache Parquet(Snappy压缩)格式存储,可通过Athena、DuckDB、PyArrow、pandas等工具实现直接原位查询。本数据集为药物发现领域提供了可供直接开展分析的化学信息学 (Cheminformatics) 基底,可应用于虚拟筛选、类药性(QED,定量类药性评分,Quantitative Estimate of Drug-likeness)过滤、相似性搜索、化学空间探索以及大规模ADMET模型训练——无需重复计算描述符与属性预测结果以跨……



