Curated Excited-State Molecular Database for Machine Learning
收藏资源简介:
This repository contains the curated benchmark dataset used in our study of interpretable machine learning for molecular excited-state prediction. The deposited CSV file contains approximately 36,000 molecular entries described by 1,041 columns, including a molecular identifier, 11 retained scalar molecular and electronic descriptors, 1024-bit Morgan fingerprint features, and five TD-DFT target properties: E(S1), f(S1), E(S2), E(T1), and E(T2). The scalar descriptors are NAts, HOMO, LUMO, TPSA, Num_H_Donors, Num_H_Acceptors, Bonds_Rotatable, Rings, Rings_Aromatic, Rings_Fused, and Hbonds. This version corresponds to the processed, model-ready dataset used in the associated manuscript and is provided to support reproducibility of benchmark development, model training, and evaluation. The original excited-state values were taken from https://www.nature.com/articles/s41597-022-01142-7; this Zenodo record provides the curated working dataset used in this study.



