insilicomedicine/URSA-benchmarking-sets
收藏资源简介:
URSA基准测试集是一个用于单步逆合成(retrosynthesis)的基准数据集,来自Zagribelnyy等人(2026年)的研究。该数据集包含三个CSV文件:URSA-expert-2026.csv(包含100个专家精选的具有挑战性的药物化学目标分子,具有高复杂性、多样化的化学类型和典型的药物分子立体化学)、USPTO-50K-test.csv(从USPTO-50K测试集中精选的4,972个产品分子,经过筛选去除了可能影响基准语义的条目)和USPTO-50K-test-mini.csv(USPTO-50K-test的10%固定随机子样本,包含497个目标,用于快速迭代和低成本评估)。每个文件提供产品分子的SMILES表示(在适用时保留立体化学),旨在用于评估逆合成模型的性能。数据集基于Apache License 2.0许可发布,部分数据源自USPTO-50K测试集。
The URSA benchmark dataset is a standardized resource for evaluating single-step retrosynthesis, originating from the 2026 study by Zagribelnyy et al. This dataset comprises three CSV files: URSA-expert-2026.csv, which contains 100 expert-curated challenging pharmaceutically relevant target molecules with high complexity, diverse chemical scaffolds, and typical stereochemistry of drug-like molecules; USPTO-50K-test.csv, which includes 4,972 product molecules selected from the USPTO-50K test set, with entries that may compromise the semantic validity of the benchmark filtered out; and USPTO-50K-test-mini.csv, a fixed 10% random subsample of the USPTO-50K-test set containing 497 target molecules for rapid iteration and low-cost evaluation. Each file provides the SMILES representations of the product molecules, with stereochemistry preserved where applicable, and is intended for evaluating the performance of retrosynthesis models. This dataset is released under the Apache License 2.0, and part of its data is derived from the USPTO-50K test set.




