Ilm-NMR-P31
收藏资源简介:
This publication introduces a novel open-access 31P Nuclear Magnetic Resonance (NMR) shift database designed to bridge the gap between commercial and open-access resources. With 15,967 entries encompassing 15,405 distinct molecules from 3,996 references, this database offers a comprehensive repository of organic and inorganic compounds. Emphasizing single-phosphorus atom compounds, the database facilitates data mining and machine learning endeavors, particularly in signal prediction and Computer-Assisted Structure Elucidation (CASE) systems. Version 2.2 adds data for a comparision of functional performance on the Density Functional Theory (DFT)-derived 31P NMR shifts. For the structures optimized in vacuum the NMR shifts were calculated using B3LYP, KT3 (_kt3) and r2SCAN (_r2s). The latter two are specialized functionals for NMR shift calculation, while B3LYP can be seen as a generalist functional. Accordingly, only the csv and the rda file were updated as no new structures nur geometries were added. Version 2.1 is a revised version of 2.0. The conformer analysis was updated to contain a geometry optimization step after the initial conformer ensemble generation. This slightly altered the associated geometries and Density Functional Theory (DFT)-derived 31P NMR shifts and shielding constants. In Version 2.0 of this dataset the following changes were made: 66 new single phosphorous containing molecules as well as 1,825 components which contain more than one phosphorous atom, but which are symmetric, so that all phosphorous atoms show the same 31P NMR shift were added. For 10,111 molecules Density Functional Theory (DFT)-derived 31P NMR shifts and shielding constants are available. This includes single point calculations in vacuum as well as in five implicit solvents (chloroform, dimethyl sulfoxide, water, toluene, acetonitrile). Furthermore, also Boltzmann-weighted NMR shifts from up to 20 conformers are available. The repository also lists the DFT-optimized xyz coordinates for each of the 10,111 molecules, including all xyz information for the up to 20 conformers considered. The xyz files are named as the structures in the dataset. The dataset is available in 3 different formats: The molecular structures are available as either a single large SDF file, which includes all database information or as multiple smaller SDF (Structure Data Format) files, where each SDF file lists the information for one molecule. The NMR data is stored in the tag section of the SDF files in the format proposed by the NMReData initiative (http://nmredata.org/, V1.1). This includes the 31P shift the solvent, where available. The data is also available as CSV file. The file does not contain the molecular structures but rather lists important information (sum formula, 31P shift, solvent, molecular weight, number of carbon, nitrogen, phosphorous and oxygen atoms) of which the most important is the canonical SMILES string generated by OpenBabel V3.1.1. The same information as in the CSV and SDF files is also available in a RDA file, which is the R Data Format for the script language R (https://www.r-project.org/). The data is stored as a "tibble" (https://tibble.tidyverse.org/) which is a special data frame format in R.
本出版物介绍了一款全新的开放获取型31P核磁共振(Nuclear Magnetic Resonance, NMR)位移数据库,旨在填补商用与开放获取资源之间的空白。该数据库包含15967条条目,涵盖来自3996篇参考文献的15405种不同分子,是一套涵盖有机与无机化合物的综合性数据仓库。数据库重点关注单磷原子化合物,可为数据挖掘与机器学习任务,尤其是信号预测及计算机辅助结构解析(Computer-Assisted Structure Elucidation, CASE)系统,提供有力支撑。 版本2.2新增了基于密度泛函理论(Density Functional Theory, DFT)计算得到的31P NMR位移的泛函性能对比数据。针对在真空环境下优化的结构,分别采用B3LYP、KT3(_kt3)以及r2SCAN(_r2s)泛函计算其NMR位移。后两者为专门用于NMR位移计算的专用泛函,而B3LYP则属于通用型泛函。由于未新增结构或几何构型,因此仅更新了CSV与RDA格式文件。 版本2.1是2.0的修订版本。其中构象分析得到更新,在初始构象集合生成后新增了几何优化步骤。这一改动轻微调整了相关几何构型以及基于密度泛函理论(DFT)计算得到的31P NMR位移与屏蔽常数。 本数据集2.0版本做出了如下更新: 66种单磷化合物以及1825种含多个磷原子但结构对称、所有磷原子均表现出相同31P NMR位移的组分被新增至数据库中。 针对10111种分子,可获取基于密度泛函理论(DFT)计算得到的31P NMR位移与屏蔽常数,包括真空环境下以及五种隐式溶剂(氯仿、二甲基亚砜、水、甲苯、乙腈)中的单点能计算结果。此外,还可获取最多20个构象的玻尔兹曼加权NMR位移数据。 该仓库还提供了10111种分子的DFT优化xyz坐标文件,包含所考虑的最多20个构象的全部xyz信息。xyz文件以数据集中对应的结构命名。 本数据集提供三种不同格式的获取方式: 分子结构可采用单个包含全部数据库信息的大型SDF(Structure Data Format)文件,或多个分别存储单分子信息的小型SDF文件。NMR数据以NMReData倡议(http://nmredata.org/,V1.1)提议的格式存储于SDF文件的标签段中,其中包含31P位移及可用的溶剂信息。 数据亦可通过CSV文件获取。该文件不包含分子结构,而是列出了关键信息:分子式、31P位移、溶剂、分子量、碳、氮、磷、氧原子的数量,其中最为重要的是由OpenBabel V3.1.1生成的标准SMILES字符串。 与CSV及SDF文件相同的信息也可通过RDA文件获取,该格式为脚本语言R(https://www.r-project.org/)的R数据格式。数据以"tibble"格式存储,这是R语言中一种特殊的数据框格式。



