遇见数据集

[Data augmentation in a TTL] - Fictive dataset (27.5M) with up to 5k reactions per template // (13'953 template extracted from USPTO-FULL IBM version)

收藏
Zenodo2025-10-09 更新2026-05-26 收录
官方服务:

资源简介:

Full generated fictive dataset, containing 27.5M reactions with up to 5000 reactions per radius 1 reaction template (13'953 reaction templates from USPTO-full, IBM version). Title of the manuscript: "Data augmentation in a Triple Transformer Loop retrosynthesis model" Abstract: Reactions in the US Patent Office (USPTO) are biased towards a few over-represented reaction types, which potentially limits its usefulness for computer-assisted synthesis planning (CASP). To obtain an equilibrated dataset, we applied retrosynthesis templates to USPTO molecules as products (P) to generate starting materials (SM). We then used transformer T2 from our recently reported triple transformer loop (TTL) retrosynthesis model to predict reagents (R) for the SM®P reaction. Finally, we validated the prediction by requesting a high confidence prediction (>95%) for the prediction of P from SM+R by TTL transformer T3. We generated up to 5,000 reactions per template, resulting in 27.5 million validated fictive reactions covering the chemical space of the original UPSTO dataset. To exemplify the use of this dataset, we show that a single-step retrosynthesis transformer model trained with a template equilibrated subset of 1,097,374 fictive reactions outperforms the corresponding model trained on USPTO reactions only.

提供机构:
Zenodo
创建时间:
2024-07-29
二维码
社区交流群
二维码
科研交流群
商业服务