InstaDeepAI/SPICE2_curated_v2
收藏资源简介:
SPICE2_curated_v2数据集是基于OMOL25数据集的SPICE子集,包含约200万个结构,这些结构在ωB97M-V/def2-TZVPD理论级别下计算。与之前版本相比,该版本更新了数据源(使用OMOL25的SPICE子集替代原始SPICE数据集),明确包含了带电系统,并通过添加钙、钾、锂、镁和钠等元素,将覆盖元素从12个扩展到17个。数据集按90/9/1的比例分割为训练集、验证集和测试集,分割基于分子SMILES,确保同一分子的不同构象不出现在不同集合中。训练集包含1,763,962个结构,覆盖17种化学元素(硼、溴、碳、钙、氯、氟、氢、碘、钾、锂、镁、氮、钠、氧、磷、硫、硅);验证集包含172,838个结构,覆盖17种元素;测试集包含17,477个结构,覆盖13种元素(这是按SMILES分割的自然结果)。过滤过程移除了非物理结构(仅保留所有氢原子恰好有一个键的结构)和高力结构(应用总力过滤器0.1 eV/Å和最大力过滤器15 eV/Å)。该数据集用于训练新的预训练模型,并基于OMOL25数据集(来源:https://doi.org/10.48550/arXiv.2505.08762)。
SPICE2_curated_v2 is a dataset based on the SPICE subset of the OMOL25 dataset, comprising approximately 2 million structures computed at the ωB97M-V/def2-TZVPD level of theory. Compared to previous versions, it introduces an updated data source (filtering on the SPICE subset of OMOL25 instead of the original SPICE dataset), explicit inclusion of charged systems, and expanded chemical coverage from 12 to 17 elements by adding Ca, K, Li, Mg, and Na. The dataset is split into training, validation, and test sets with a 90/9/1 ratio, partitioned per molecular SMILES to ensure different conformations of the same molecule do not appear across sets. The training set contains 1,763,962 structures covering 17 chemical elements (B, Br, C, Ca, Cl, F, H, I, K, Li, Mg, N, Na, O, P, S, Si), the validation set has 172,838 structures across 17 elements, and the test set contains 17,477 structures across 13 elements (a natural result of the per-SMILES split). The filtering process removed unphysical structures (keeping only those where all hydrogen atoms have exactly one bond) and high-force structures (applying a total force filter of 0.1 eV/Å and a maximum force filter of 15 eV/Å). It is used to train new pre-trained models and is sourced from the OMOL25 dataset (https://doi.org/10.48550/arXiv.2505.08762).




