Supporting data for "Multi-head committees enable direct uncertainty prediction for atomistic foundation models"
收藏资源简介:
The data is sorted by the respective datasets that were used to train and analyze the models. These datasets for custom-made models are 3BPA, water, and rMD17. There is a separate directory for the foundation model. The directories contain the training datasets, the scripts used to create the training sets, training scripts, and the trained models. 3BPA: Contains directories for different training set sizes. Most work was done on the 100-strucutre training set contained in the `trainset_100` directory. This directory also contains sub-directories that were used for work in the supplementary material, such as committees without training set subsampling. water: Contains the training, validation, and test sets for the analysis of bulk liquid water. rMD17: The directory `full_trainset` contains the material used for the work found in the main manuscript. Additionally, there are directories `wo_*`, which contain the material for the models where one molecule was excluded. The analysis of these models can be found in the supplementary material. foundation: The most important directory of this part is `qbc/dataset_8000`, containing the QbC-selected reduced training dataset of MPtrj and a trained multihead committee model with and without the original pre-trained head. Additionally, there are models trained on only the 1000 structures selected first by the QbC as well as models trained on randomly selected structures and the structures in the MPtrj dataset with the maximum mean force.
本数据集按照用于模型训练与分析的各类数据集进行分类整理。本次用于定制化模型训练的数据集包括3BPA、water以及rMD17,另有单独的目录用于存放基础模型(foundation model)。各目录均包含训练数据集、用于构建训练集的脚本、训练脚本以及已训练完成的模型。 3BPA:该目录下包含对应不同训练集规模的子目录。绝大多数实验基于`trainset_100`目录下包含100个结构的训练集开展,该目录同时包含用于补充材料相关研究的子目录,例如未对训练集进行二次采样的委员会模型相关实验目录。 water:该目录包含用于体相液态水分析的训练集、验证集与测试集。 rMD17:`full_trainset`目录存放了主论文中相关研究所使用的数据集。此外,`wo_*`系列目录对应剔除了单种分子的模型所用数据集,针对这些模型的分析结果见于补充材料。 foundation:该部分目录中最为关键的是`qbc/dataset_8000`,其中包含针对MPtrj数据集经QbC(Query-by-Committee)筛选得到的精简训练数据集,以及带有与不带有原始预训练头的已训练多头委员会模型。此外还包含仅使用QbC首轮筛选出的1000个结构训练的模型、仅使用随机筛选结构训练的模型,以及MPtrj数据集中平均力最大的结构对应的模型。



