completeRXN-benchmark-26/completeRXN
收藏资源简介:
CompleteRxn Benchmark是一个用于反应完成任务的基准测试数据集,包含206,423个反应。任务是根据一个原子不平衡的USPTO反应SMILES,预测缺失的分子以产生平衡的方程式。数据集提供了三种分割类型(随机、组OOD、极端OOD),每种类型有5次重复,以测试不同结构新颖性水平的泛化能力。数据集的特征包括唯一标识符、原始反应SMILES、输入反应SMILES、目标反应SMILES等。数据集的来源包括USPTO和FlowER的数据,并提供了详细的列描述和任务指标。
The CompleteRxn Benchmark is a dataset for benchmarking the task of reaction completion, containing 206,423 reactions. The task is to predict the missing molecules to produce a balanced equation given an atom-unbalanced USPTO reaction SMILES. The dataset provides three split types (random, group OOD, extreme OOD) with 5 repetitions each to test generalization across structural novelty levels. The features of the dataset include unique identifiers, original reaction SMILES, input reaction SMILES, target reaction SMILES, etc. The dataset sources include data from USPTO and FlowER, and provides detailed column descriptions and task metrics.




