遇见数据集

Ilm-NMR-P31

收藏
Zenodo2025-11-19 更新2026-05-26 收录
官方服务:

资源简介:

This publication introduces a novel open-access 31P Nuclear Magnetic Resonance (NMR) shift database designed to bridge the gap between commercial and open-access resources. With 15,967 entries encompassing 15,405 distinct molecules from 3,996 references, this database offers a comprehensive repository of organic and inorganic compounds. Emphasizing single-phosphorus atom compounds, the database facilitates data mining and machine learning endeavors, particularly in signal prediction and Computer-Assisted Structure Elucidation (CASE) systems. Version 2.1 is a revised version of 2.0. The conformer analysis was updated to contain a geometry optimization step after the initial conformer ensemble generation. This slightly altered the associated geometries and Density Functional Theory (DFT)-derived 31P NMR shifts and shielding constants. In Version 2.0 of this dataset the following changes were made: 66 new single phosphorous containing molecules as well as 1,825 components which contain more than one phosphorous atom, but which are symmetric, so that all phosphorous atoms show the same 31P NMR shift were added. For 10,111 molecules Density Functional Theory (DFT)-derived 31P NMR shifts and shielding constants are available. This includes single point calculations in vacuum as well as in five implicit solvents (chloroform, dimethyl sulfoxide, water, toluene, acetonitrile). Furthermore, also Boltzmann-weighted NMR shifts from up to 20 conformers are available. The repository also lists the DFT-optimized xyz coordinates for each of the 10,111 molecules, including all xyz information for the up to 20 conformers considered. The xyz files are named as the structures in the dataset. The dataset is available in 3 different formats: The molecular structures are available as either a single large SDF file, which includes all database information or as multiple smaller SDF (Structure Data Format) files, where each SDF file lists the information for one molecule. The NMR data is stored in the tag section of the SDF files in the format proposed by the NMReData initiative (http://nmredata.org/, V1.1). This includes the 31P shift the solvent, where available. The data is also available as CSV file. The file does not contain the molecular structures but rather lists important information (sum formula, 31P shift, solvent, molecular weight, number of carbon, nitrogen, phosphorous and oxygen atoms) of which the most important is the canonical SMILES string generated by OpenBabel V3.1.1. The same information as in the CSV and SDF files is also available in a RDA file, which is the R Data Format for the script language R (https://www.r-project.org/). The data is stored as a "tibble" (https://tibble.tidyverse.org/) which is a special data frame format in R.

提供机构:
Zenodo
创建时间:
2025-11-19
二维码
社区交流群
二维码
科研交流群
商业服务