PolySet: Statistical Ensemble Polymer Embeddings Dataset for Machine Learning
收藏资源简介:
PolySet Dataset and Training Script This archive contains the dataset and the training script used to reproduce the machine-learning results presented in the manuscript: “PolySet: Restoring the Statistical Ensemble Nature of Polymers for Machine Learning”K. Ferji (2025) Contents 1. PolySet_dataset.csv A curated dataset of 10,000 synthetic homopolymers used in the PolySet study.Each record includes: Mn – number-average molar mass D - Dispersity Mz – molar-mass moment z Mz1 – molar-mass moment z+1 (prediction target) Fp – PolySet distribution-aware embedding (JSON list) The embeddings are fully precomputed, allowing users to evaluate representation effects without recomputing chain ensembles. 2. training_Mz1.py A minimal and fully reproducible PyTorch script that trains a lightweight neural regressor on the PolySet embeddings to predict Mz+1.It reproduces the learning curves and accuracy values reported in the manuscript.



