遇见数据集

SMPI Dataset: Molecular Descriptors and SMILES for 33,715 Organic Molecules (smpi-nog33715mols2)

收藏
Zenodo2026-09-26 更新2026-10-01 收录
官方服务:

资源简介:

This dataset contains structural molecular property indices (SMPI) calculated for a benchmark collection of 33,715 organic molecules. The dataset is provided as a tab-separated text file (smpi-nog33715mols2.txt) designed for chemoinformatics analysis, quantitative structure-activity/property relationship (QSAR/QSPR) modeling, and machine learning workflows. Note on data hosting: This repository serves as the official open-access host for the dataset smpi-nog33715mols2, superseding the legacy HTTP server location (http://193.226.7.140/~lori/data/smpi-nog233715mols2.tgz). Dataset Contents & Structure File Name: smpi-nog33715mols2.txt Format: Tab-Separated Values (TSV / UTF-8 text) Number of Records: 33,715 chemical structures (plus 1 header line) Number of Columns: 125 columns (1 SMILES string + 124 SMPI descriptors) Data Fields & Column Nomenclature The first column contains the molecular structure represented as a 1D canonical SMILES string. The subsequent 124 columns contain topological and structural molecular property indices (SMPI) derived from chemical graph operators, molecular matrices, and atomic property weightings. Column Encoding Standard: Each descriptor column name (e.g., RNEUA, IJETB, LEUTB, LNPUG) follows a systematic 5-to-6-character naming convention: Operator / Function Prefix (1–2 letters): R: Reciprocal / Ratio operator L: Logarithmic transformation I: Integral / Inverse sum operator J: Wiener / Hosoya-type distance function M: Maximum / Mean metric evaluation N: Normalized vertex/edge structural count F: Fragmental / Fourier graph transform component Weighting / Property Code (1–2 letters): E: Electronegativity weighting P: Atomic polarizability / Partial charge weighting U: Unweighted (pure topological distance / degree metric) D: Topological / Geometric distance weighting Graph Matrix & Polynomial Variant Suffix (A through G): Denotes the specific molecular matrix operator family (e.g., Szeged, Cluj, Distance-Extended, Adjacency) and polynomial evaluation root/point. Applications QSAR/QSPR Modeling: High-throughput feature selection and regression for bioactivity and physical property prediction. Machine Learning & Deep Learning: Benchmark training set for graph neural networks, random forests, and gradient boosting approaches in computational chemistry. Molecular Diversity & Similarity Search: Chemical space mapping based on multi-dimensional topological invariants. Usage Example (Python / Pandas) import pandas as pd # Load dataset into Pandas DataFrame df = pd.read_csv('smpi-nog33715mols2.txt', sep='\t') print(f"Dataset shape: {df.shape}") print(df.head())

提供机构:
Zenodo
创建时间:
2026-09-26
二维码
社区交流群
二维码
科研交流群
商业服务