遇见数据集

modelforge curated dataset: tmQM-xtb

收藏
Zenodo2025-05-12 更新2026-05-26 收录
官方服务:

资源简介:

Curated tmQM-xtb Dataset: - T=300K dataset restricted to [Pd, Zn, Fe, and Cu]- Version: v1.1_PdZnFeCu_T300K This dataset contains 23177 unique systems with 490,861 total configurations, sampled at T=300K. This dataset is limited to systems that contain transition metals Pd, Zn, Fe, or Cu, and also only contain elements C, H, P, S, O, N, F, Cl, or Br. Potentially problematic configurations (i.e., unstable or those with underlying structural changes) were removed. Briefly, bond inference was performed on the initial configuration using RDKit and a configuration was excluded if any of those bond distances changed by more than 0.15 angstroms compared to the initial, energy minimized state. This dataset was generated starting from the tmQM dataset; the original tmQM repository (https://github.com/uiocompcat/tmQM) was forked and a release made that corresponds to the data committed on 13 August 2024 (https://github.com/chrisiacovella/tmQM/releases/tag/2024Aug13). Each molecule was evaluated using gfn2-xtb, and then a short MD simulation performed to provide additional configurations of the molecules. The tblite package was used to evaluate the energetic of the system using the gfn2-xtb formalism. MD simulations were performed using the Atomic Simulation Environment (ASE), using the Langevin integrator Simulations were performed at 300K with a 1 fs timestep and 0.01 1/fs friction damping factor. In all trajectories, the first configuration corresponds to the energy minimized configuration reported in the original tmQM dataset. 100 steps were taken between snapshots, with 30 total snapshots per molecule During MD sampling, gfn2-xtb accuracy was set to 2; all reported properties were calculated at accuracy level 1. Scripts used to perform the sampling can be found https://github.com/chrisiacovella/xtb_config_gen Properties included: atomic_numbers positions "per_atom" "nanometer" forces "per_atom" "kilojoule_per_mole / nanometer" partial_charges "per_atom" "elementary_charge" energies "per_system" "kilojoule_per_mole" dipole_moment_per_system "per_system" "elementary_charge * nanometer" total_charge "per_system" "elementary_charge" spin_multiplicities "per_system" "dimensionless" stoichiometry "meta_data"

精选tmQM-xtb数据集:T=300K限定于[钯(Pd)、锌(Zn)、铁(Fe)、铜(Cu)]的数据集 版本:v1.1_PdZnFeCu_T300K 本数据集包含23177个独特体系,总计490861个构型,采样温度设定为300K。 本数据集仅包含含过渡金属钯、锌、铁或铜的体系,且体系仅由碳(C)、氢(H)、磷(P)、硫(S)、氧(O)、氮(N)、氟(F)、氯(Cl)或溴(Br)元素构成。 已剔除存在潜在问题的构型(即不稳定或发生潜在结构变化的构型)。简言之,使用RDKit对初始构型进行键长推断,若某构型的任意键长与初始能量最小化状态的差值超过0.15埃,则将该构型剔除。 本数据集源自tmQM数据集;我们复刻了原始tmQM仓库(https://github.com/uiocompcat/tmQM),并基于2024年8月13日提交的数据发布了对应版本(https://github.com/chrisiacovella/tmQM/releases/tag/2024Aug13)。 首先使用gfn2-xtb对每个分子进行能量与结构评估,随后执行短时分子动力学(MD)模拟以生成该分子的额外构型。 使用tblite软件包结合gfn2-xtb形式化方法计算体系的能量。 分子动力学模拟采用原子模拟环境(Atomic Simulation Environment, ASE),使用朗之万积分器。 模拟在300K温度下进行,时间步长为1飞秒(fs),摩擦阻尼因子为0.01 1/fs。 所有轨迹的首个构型均与原始tmQM数据集中报告的能量最小化构型一致。 快照采样间隔为100步,每个分子总计生成30个快照。 在分子动力学采样过程中,gfn2-xtb的计算精度设为2;所有上报的属性均在精度等级1下完成计算。 用于执行采样的脚本可在https://github.com/chrisiacovella/xtb_config_gen获取。 包含的属性如下: - 原子序数(atomic_numbers):逐原子维度 - 原子坐标(positions):逐原子(per_atom),单位:纳米(nanometer) - 原子受力(forces):逐原子(per_atom),单位:千焦每摩尔每纳米(kilojoule_per_mole / nanometer) - 部分电荷(partial_charges):逐原子(per_atom),单位:元电荷(elementary_charge) - 体系能量(energies):逐体系(per_system),单位:千焦每摩尔(kilojoule_per_mole) - 体系偶极矩(dipole_moment_per_system):逐体系(per_system),单位:元电荷·纳米(elementary_charge * nanometer) - 体系总电荷(total_charge):逐体系(per_system),单位:元电荷(elementary_charge) - 自旋多重度(spin_multiplicities):逐体系(per_system),无量纲(dimensionless) - 化学计量比(stoichiometry):元数据(meta_data)

提供机构:
Zenodo
创建时间:
2025-04-16
二维码
社区交流群
二维码
科研交流群
商业服务