Assay2Mol datasets
收藏资源简介:
This dataset contains supplementary files for the manuscript "Assay2Mol: Large Language Model-based Drug Design Using BioAssay Context". BioAssay embedding The BioAssay embedding vectorbase is embedded with the OpenAI text-embedding-ada-002 model. We download BioAssay json data from the PubChem FTP site. If you use the embedding vectorbase, please see the PubChem download policies and citation guidelines. The vectorbase is stored with Faiss in the files index.pkl and index.faiss. The dataset is current as of September 28, 2024. CrossDocked2020 test set We provide the CrossDocked2020 test set as test_set.zip with the processed pqr and pdbqt formatted protein structures, making it more convenient for docking. The files are processed with the script from TargetDiff. Please extract the zip file contents to the docking/test_set/ directory of the Assay2Mol GitHub repository. See the CrossDocked2020 dataset and manuscript for more information.



