merged_cinn_dataset.csv: A Merged Thermodynamic Dataset of Sigma Profiles, Quantum Indicators, and Activity Coefficients for Training and Validation of KANA AI
收藏资源简介:
Overview This dataset, merged_cinn_dataset.csv, is a consolidated thermodynamic dataset comprising over 287,000 data entries compiled for the training and validation of KANA AI, a deep learning model designed to predict activity coefficients of molecular mixtures. The dataset integrates molecular sigma profiles, quantum chemical indicators, and experimental or computed activity coefficient values (expressed as ln γ) across a diverse range of chemical systems and thermodynamic conditions. Dataset contents Sigma profiles — Each entry includes a discretised sigma profile derived from quantum chemical COSMO (Conductor-like Screening Model) calculations. The sigma profile encodes the surface charge density distribution of a molecule, serving as the primary molecular descriptor for capturing intermolecular interactions in liquid-phase systems within the COSMO-SAC framework. Quantum indicators — Supplementary quantum chemical descriptors are provided alongside each sigma profile. These include molecular volume, accessible surface area, dipole moment, and frontier molecular orbital energies (HOMO and LUMO). These indicators enrich the molecular representation beyond the sigma profile and improve the model's capacity to generalise across structurally diverse compounds. Activity coefficients (ln γ) — The natural logarithm of the activity coefficient serves as the primary regression target. Activity coefficients quantify the degree of thermodynamic non-ideality in liquid-phase mixtures and are a critical property in phase equilibrium calculations, separation process design, and chemical process modelling. Intended use The dataset is partitioned for two purposes. The training split is used to optimise the parameters of the KANA AI neural network architecture, enabling it to learn the mapping from sigma profiles and quantum indicators to ln γ values. The validation split is reserved for assessing the generalisation capability of the trained model on molecular systems and thermodynamic conditions not encountered during training. Applicability This dataset is intended for use in machine learning–based thermodynamic property prediction, with particular relevance to activity coefficient modelling, COSMO-SAC model development, and data-driven approaches in chemical engineering. Researchers working on predictive thermodynamics, molecular property estimation, or neural network-based equation-of-state models may find this dataset directly applicable.



