DFT Calculated xyz and log Files as well as csv Files for Machine Learning in Support of "Tailoring Phosphine Ligands for Improved C H Activation: Insights from Δ-Machine Learning"
收藏资源简介:
Transition metal complexes have played crucial roles in various homogeneous catalytic processes due to their exceptional versatility. This adaptability stems not only from the central metal ions but also from the vast array of choices of the ligand spheres, which form an enormously large chemical space. For example, Rh complexes, with a well-designed ligand sphere, are known to be efficient in catalyzing the C-H activation process in alkanes. To investigate the structure-property relation of the Rh complex and identify the optimal ligand that minimizes the calculated reaction energy ΔE of an alkane C-H activation, we have applied a Δ-Machine Learning method trained on various features to study 1,743 pairs of reactants (Rh(PLP)(Cl)(CO)) and intermediates (Rh(PLP)(Cl)(CO)(H)(propyl)). Our findings demonstrate that the models exhibit robust predictive performance when trained on features derived from electron density (R2 = 0.816), and SOAPs (R2 = 0.819), a set of position-based descriptors. Leveraging the model trained on xTB-SOAPs that only depend on the xTB-equilibrium structures, we propose an efficient and accurate screening procedure to explore the extensive chemical space of bisphosphine ligands. By applying this screening procedure, we identify ten newly selected reactant-intermediate pairs with an average ΔE of 33.2 kJ mol-1, remarkably lower than the average ΔE of the original data set of 68.0 kJ mol-1. This underscores the efficacy of our screening procedure in pinpointing structures with significantly lower energy levels. _______________________________________________________________________ The dataset contains three file types: Version 1.0: xyz files of the final optimized Rh-phosphine complexes; one set for the starting materials denoted as "molecule-XXXX_4-times" and one set for the intermediates after C-H activation denoted as "molecule-XXXX_6-times" Gaussian16 log files for the optimization process; one set for the starting materials denoted as "molecule-XXXX_4-times" and one set for the intermediates after C-H activation denoted as "molecule-XXXX_6-times" csv files containing the per molecule features used for training the different machine learning models. The name of the csv files indicates which property was predicted and which model was used New in version 1.1 (other data is unchanged): Gaussian16 log files for the ten newly identified bisphosphine ligands; one set for the product material denoted as "LXX_6-times-axial" and one set for the transition state for the C-H activation denoted as "LXX_C-H-activation_TS"
过渡金属配合物凭借其卓越的多功能性,在各类均相催化过程中发挥了至关重要的作用。这种适应性不仅源自中心金属离子,还来自配体球(ligand spheres)丰富多样的选择,由此构建出规模极为庞大的化学空间。举例而言,经过合理配体球设计的铑(Rh)配合物,已被证实可高效催化烷烃的C-H活化过程。为探究铑配合物的结构-性能关系,并筛选出能使烷烃C-H活化计算反应能量ΔE降至最低的最优配体,我们采用了基于多种特征训练的Δ机器学习(Δ-Machine Learning)方法,对1743对反应物与中间体进行了研究:反应物为Rh(PLP)(Cl)(CO),中间体为Rh(PLP)(Cl)(CO)(H)(丙基)。 研究结果表明,当模型基于电子密度衍生特征(决定系数R²=0.816)以及基于位置的描述符SOAPs(R²=0.819)进行训练时,展现出极佳的预测性能。我们借助仅依赖xTB平衡结构的xTB-SOAPs训练得到的模型,提出了一套高效且精准的筛选流程,用于探索双膦配体的广阔化学空间。通过该筛选流程,我们筛选出10组新的反应物-中间体对,其平均ΔE为33.2 kJ·mol⁻¹,显著低于原始数据集的平均ΔE(68.0 kJ·mol⁻¹),这充分证明了我们的筛选流程可有效定位具有更低反应能量的结构。 _______________________________________________________________________ 本数据集包含三类文件: 版本1.0: 1. 最终优化后的铑-膦配合物的xyz文件:一组为反应物(记为"molecule-XXXX_4-times"),另一组为C-H活化后的中间体(记为"molecule-XXXX_6-times") 2. 优化过程的Gaussian16日志文件:同样分为反应物("molecule-XXXX_4-times")与C-H活化中间体("molecule-XXXX_6-times")两类 3. 用于训练不同机器学习模型的单分子特征csv文件:文件名可指示所预测的物性与所用的模型类型 版本1.1新增内容(其余数据保持不变): 针对10组新筛选出的双膦配体的Gaussian16日志文件:一组为产物(记为"LXX_6-times-axial"),另一组为C-H活化过程的过渡态(记为"LXX_C-H-activation_TS")



