遇见数据集

Training and Test Splits of datasets-Building informative materials datasets beyond targeted objectives

收藏
Zenodo2026-05-26 更新2026-05-29 收录
官方服务:

资源简介:

This repository contains the training-validation and test datasets used in the study “Building informative materials datasets beyond targeted objectives.” The study used four DFT datasets and one experimental dataset. For each case, 10% of the data was held out as a test set, while the remaining 90% was used as the candidate pool for dataset construction. This repository provides these splits to make clear which data were used for dataset construction and which data were used for model evaluation. The training-validation datasets correspond to the candidate pools used to build the datasets. The test datasets correspond to the hold-out sets used to evaluate the performance of models trained on the constructed datasets. For each dataset, the feature matrix and outcome matrix are provided separately. In the DFT datasets, the outcome labels y1, y2, and y3 correspond to formation energy, bulk modulus, and bandgap, respectively. Datasets labeled RAW correspond to the original DFT datasets before data cleaning. Datasets labeled All, but not RAW, correspond to the filtered datasets used in the paper. For the experimental dataset, the training-validation and test splits are provided from the curated dataset used in the study. All rows keep their original indices. These indices can be used to track which samples were selected during dataset construction in each case.

提供机构:
Zenodo
创建时间:
2026-05-26
二维码
社区交流群
二维码
科研交流群
商业服务