CAT: A Data-Driven Framework for Solvent Selection and Optimisation
收藏资源简介:
This Zenodo repository contains the data used to train our model CAT, presented in CAT: A Data-Driven Framework for Solvent Selection and Optimisation. A reaction database consisting of reactions retrieved from a patent database, including a large set of atom-mapped SMILES reactions, their corresponding ID from patent literature and the reaction class in which they originate from, available at Reaction Dataset.xls. The reaction database including a set of featurisations describing the reactants using the Mordred featuriation module at RDKit, available at Featurised Reactants.csv. A dataset consisiting of reactions sampled from the orginal dataset, used to train our model, available at Sampled Reactions inc Featurisations.csv & Sampled Reactions wv Rate Constant & Featuriation.csv. An additional dataset used to trained the model, including a set of Menschutkin reactions, available at Sampled Reactions wv Rate constant & Featurisation wv Menschutkin Rxn.xls. A Minnesota Solvent Database, containg quantumn solvent descriptors determined by a universal solvation model, orginally compiled by Paul Winget & Co-authors, cleaned by filling in missing values using a self-built random forest regressor and additionally removing two solvents (ethanol and water), available at Cleaned Minnesota Solvent Database.xls.



