Compound selection procedures based on molecular similarity and diversity are widely used in drug discovery. Current algorithms are often time consuming when applied to very large compound sets. This
THEOBROMA is an aggregated open natural-products database containing 1,133,004 compounds from 29 source databases across six biological kingdoms. The corpus preserves full-stereochemistry InChIKeys an
Among all 270,540 drug-protein pairs from the ChEMBL data set, the top 50 unknown pairs determined by the KL1LR method using data sets were checked, and the unknown pair was listed if it was found in