遇见数据集

PI1M: A Benchmark Database for Polymer Informatics

收藏
Figshare2020-06-15 更新2026-04-08 收录
官方服务:

资源简介:

Open source data in large scale are the cornerstones for data-driven research, but they are not readily available for polymers. In this work, we build a benchmark database, called PI1M (referring to ~1 million polymers for polymer informatics), to provide data resources that can be used for machine learning research in polymer informatics. A generative model is trained on ~12,000 polymers manually collected from the largest existing polymer database PolyInfo, and then the model is used to generate ~1 million polymers. A new representation for polymers, polymer embedding (PE), is introduced, which is then used to perform several polymer informatics regression tasks for density, glass transition temperature, melting temperature and dielectric constants. By comparing the PE trained by the PolyInfo data and that by the PI1M data, we conclude that the PI1M database covers similar chemical space as PolyInfo, but significantly populate regions where PolyInfo data are sparse. We believe PI1M will serve as a good benchmark database for future research in polymer informatics.

大规模开源数据是数据驱动研究的基石,但聚合物领域暂无便捷易得的此类数据。本研究构建了一款名为PI1M(指代用于聚合物信息学的约100万条聚合物数据)的基准数据库,旨在为聚合物信息学领域的机器学习研究提供可用的数据资源。研究人员从现有规模最大的聚合物数据库PolyInfo中手动收集了约12000条聚合物数据,以此训练一款生成式模型,随后利用该模型生成了约100万条聚合物数据。本文提出了一种全新的聚合物表征方法——聚合物嵌入(polymer embedding, PE),并基于该表征方法开展了多项聚合物信息学回归任务,涵盖密度、玻璃化转变温度、熔融温度与介电常数。通过对比基于PolyInfo数据训练得到的PE与基于PI1M数据训练得到的PE,本研究证实PI1M数据库与PolyInfo覆盖的化学空间相似,但显著填充了PolyInfo数据较为稀疏的区域。我们认为,PI1M将成为未来聚合物信息学领域研究的优质基准数据库。

提供机构:
Tengfei Luo; RUIMIN MA
创建时间:
2020-06-15
二维码
社区交流群
二维码
科研交流群
商业服务