Tautomer datasets used for studying the tautomer influence on cheminformatics processing and QSAR/QSPR modeling
收藏资源简介:
The presented dataset is a supplementary material for the publication: Nikolay T. Kochev, Vesselina H. Paskaleva, Nina G. Jeliazkova, Tautomerism influence on QSAR/QSPR modeling of ecotoxicity and physicochemical properties of chemical compounds Here are presented files used as input or are results achieved during process of exploring the influence of tautomeric forms over building, applying and pre-steps in QSAR/QSPR modelling. All structures used in this study were pre-processed using ChemAxon Standardizer version 5.12.2 including extraction of SMILES linear notation from sdf files, kekulization of aromatic structures, conversion of explicit hydrogen atoms to implicit ones and removal of stereo information. The tautomeric forms for the structures were generated by means of Ambit-Tautomer software [https://doi.org/10.1002/minf.201200133], applying the IA-DFS algorithm (incremental approach based on depth-first search) with tautomeric rules for 1.3 and 1.5 hydrogen shifts and removal of topologically equivalent atoms and allene atom. Tautomeric influence over fingerprints calculation: Example chemical molecule: methimazole Generated tautomeric forms: file methimazole_tautomers.smi Results uploaded for MACCS Fingerprint calculated for the tautomeric forms of methimazole using PaDEL-Descriptor software version 2.17: file methimazole-MACCS.csv Tautomeric influence over applying QSAR/QSPR models: Chosen example software: PaDEL-Descriptor software version 2.17 Chosen QSPR model implemented az molecular propery descriptor in PaDEL-Descriptors: CrippenLogP File containing all generated tautomeric forms: Crippen-LogP_all-tautomers-PaDEL.csv First column: T, contains the number of the molecular structure for which the tautomers are calculated. If there are rows with indentical numer, they are tautomeric forms of one structure. Second column: Smiles, contains the SMILES linear notations for all molecular structures and their generated tautomeric forms. Third column: E, the rank of every tautomer calculated with Ambit-Tautomer Forth column: CrippenLogP, contains the calculated results from PaDEL-Descriptors Files LogP-tauts-Padel-DESC-part01.csv to LogP-tauts-Padel-DESC-part04.csv contained the calculated 0D, 1D and 2D molecular descriptors (including CrippenLogp) from PaDEL-Descriptors. Tautomeric influence over building QSAR/QSPR models: *AMES mutagenicity model AMES mutagenicity model is built using Random Forest machine learning method with 80 trees implemented in WEKA software version 3.7.9. The molecular structure information is coded using fingerprints implemented in PaDEL_Descriptor software and post-calculation modification was applied - 623 fingerprint bits from 2212 calculated. The build model does not include any tautomeric information. The model files is given in: file AMES.model *Tetrahymena pyriformis model Archive name: TetrahymenaPyriformis_model.zip Two types of models were built: conventional (not including tautomers) and a model over weighted descriptors. Models were built over WEKA version 3.7.9 using machine learning algorithm ExtraTree and method for descriptor selection CfgSubsetEval with Best First search aproach. Training set consist of 644 structures and the extrnal validation set consist of 110 molecular structures taken from https://doi.org/10.1016/j.chemosphere.2010.11.043 The two type of models were build over PaDEL-Descriptors 0D, 1D and 2D Training and validation sets for the conventional model: trainset_CM.arff, ext.val.set_CM.arff from the archive Training and validation sets for the weighted model: trainset_WM.arff, ext.val.set_WM.arff from the archive Generated tautomers for the training set: trainset_TetrahymenaPyriformis_tautomers.xls from the archive Generated tautomers for the external validation set: ext_validations_et_TetrahymenaPyriformis_tautomers.xls from the archive Statistical results achieved for the convetional and for the weighted model: file modelling results.xls from the archive



