遇见数据集

Ilm-NMR-P31

收藏
Zenodo2026-06-19 更新2026-05-26 收录
官方服务:

资源简介:

This publication introduces a novel open-access 31P Nuclear Magnetic Resonance (NMR) shift database designed to bridge the gap between commercial and open-access resources. With 43,317 entries encompassing 39,107 distinct molecules from 6,853 references, this database offers a comprehensive repository of organic and inorganic compounds. Emphasizing symmetric phosphorus compounds, the database facilitates data mining and machine learning endeavors, particularly in signal prediction and Computer-Assisted Structure Elucidation (CASE) systems. ------------------------------------------------------------------------------------------------------- In Version 3.0 the following changes were made: 24,296 molecules were added to the dataset derived from the NMRexp dataset (v2, 10.5281/zenodo.17296666). Only single and symmetric single phosphorous compounds were copied. In addition the NMRexp data was curated by using a HOSE-code based prediction model to discard entries, which values were outside a tolerance windows of 5 ppm between reported and predicted value. The correspdonding SDF files were generated from the reported SMILES strings found in NMRexp using OpenBabel 3.1.1 . The "origin" column was removed due to it being redunadant to the "ref" column. The "freq" column is now avialable stating the recording spectrometers frequency if avialable. For the SDF files an NMReDATA assigment tag (<NMREDATA_ASSIGNMENT>) was added. This links the phosphorous atoms to the respective 31P shift values. Consequently the NMReDATA tag <NMREDATA_1D_31P> now also lists an atom labeling and potentially more than one 31P NMR value as multiple atoms with the same shift value might be part of the same molecule. If avialable the NMReDATA tags in the SDF files now also report the solvent (<NMREDATA_SOLVENT>) as well as the spectrometer frequency (found in <NMREDATA_1D_31P>). This information was already avialable in the .csv and .rda files. ------------------------------------------------------------------------------------------------------- Version 2.2 adds data for a comparision of functional performance on the Density Functional Theory (DFT)-derived 31P NMR shifts. For the structures optimized in vacuum the NMR shifts were calculated using B3LYP, KT3 (_kt3) and r2SCAN (_r2s). The latter two are specialized functionals for NMR shift calculation, while B3LYP can be seen as a generalist functional. Accordingly, only the csv and the rda file were updated as no new structures nur geometries were added. Version 2.1 is a revised version of 2.0. The conformer analysis was updated to contain a geometry optimization step after the initial conformer ensemble generation. This slightly altered the associated geometries and Density Functional Theory (DFT)-derived 31P NMR shifts and shielding constants. In Version 2.0 of this dataset the following changes were made: 66 new single phosphorous containing molecules as well as 1,825 components which contain more than one phosphorous atom, but which are symmetric, so that all phosphorous atoms show the same 31P NMR shift were added. For 10,111 molecules Density Functional Theory (DFT)-derived 31P NMR shifts and shielding constants are available. This includes single point calculations in vacuum as well as in five implicit solvents (chloroform, dimethyl sulfoxide, water, toluene, acetonitrile). Furthermore, also Boltzmann-weighted NMR shifts from up to 20 conformers are available. The repository also lists the DFT-optimized xyz coordinates for each of the 10,111 molecules, including all xyz information for the up to 20 conformers considered. The xyz files are named as the structures in the dataset. ------------------------------------------------------------------------------------------------------- The dataset is available in 3 different formats: The molecular structures are available as either a single large SDF file, which includes all database information or as multiple smaller SDF (Structure Data Format) files, where each SDF file lists the information for one molecule. The NMR data is stored in the tag section of the SDF files in the format proposed by the NMReData initiative (http://nmredata.org/, V1.1). This includes the 31P shift, the solvent, and the spectrometer frequency where available. The data is also available as CSV file. The file does not contain the molecular structures but rather lists important information (sum formula, 31P shift, solvent, molecular weight, number of carbon, nitrogen, phosphorous and oxygen atoms) of which the most important is the canonical SMILES string generated by OpenBabel V3.1.1. The same information as in the CSV and SDF files is also available in a RDA file, which is the R Data Format for the script language R (https://www.r-project.org/). The data is stored as a "tibble" (https://tibble.tidyverse.org/) which is a special data frame format in R.

提供机构:
Zenodo
创建时间:
2023-08-18
二维码
社区交流群
二维码
科研交流群
商业服务