This project contains the dataset used in the manuscript "Solving an inverse problem with generative models". The code and documentation for using this data set are available in the manuscript. See th
This directory contains sets of molecules used to train chemical language models in the paper, "Learning generative models of molecules from limited training examples." Between 1,000 and 500,000 mol
Training data and model checkpoints accompanying paper on "Automated patent extraction powers generative modeling in focused chemical spaces". If you use this data, please cite the following manuscrip