insilicomedicine/CREED
收藏资源简介:
CREED(全面反应物穷举枚举数据集)是一个大规模、质量控制的化学反应资源,专门用于单步逆合成研究。该数据集将目标产物结构与多种合理的反应物集合(单步断开)进行配对,反映了合成规划本质上是多解决方案的,而非单一标准答案。它包含两个并行版本:CREED(不含每个反应的ChemCensor评分)和CREED-CCV(包含ChemCensor评分)。数据以Parquet分片格式分发,每条记录包括一个产物SMILES和有序的反应候选列表,其中每个候选包含产物异构SMILES、反应物异构SMILES列表和可选的ChemCensor评分。数据集统计显示,CREED版本涵盖约149万产物和2271万反应候选,平均每个产物约15个候选;CREED-CCV版本涵盖约69万产物和637万反应候选,平均每个产物约9个候选。
CREED (Comprehensive Reactant Exhaustive Enumeration Dataset) is a large-scale, quality-controlled resource of chemical reactions for single-step retrosynthesis. It pairs target product structures with many plausible reactant sets (one-step disconnections), reflecting that synthesis planning is inherently multi-solution rather than single–ground-truth. The dataset includes two parallel releases: CREED (without a per-reaction ChemCensor score) and CREED-CCV (with ChemCensor scores). Data is distributed as Parquet shards, with each record consisting of a product SMILES and an ordered list of reaction candidates, where each candidate contains product isomeric SMILES, a list of reactants isomeric SMILES, and an optional ChemCensor score. Statistics show that the CREED release covers approximately 1.49 million products and 22.71 million reaction candidates, with an average of about 15 candidates per product; CREED-CCV covers approximately 698,765 products and 6.37 million reaction candidates, with an average of about 9 candidates per product.



