GoldDIGR
收藏资源简介:
GoldDIGR是一个大规模开放的DFT反应计算数据集,包含约1,000,000个xTB/DFT反应计算。数据来源于已发表化学论文的支持信息,并通过统一的处理流程(包括几何结构/过渡态搜索、IRC、单点DFT计算以及xTB自旋/结构扫描)重新计算得到。每个计算是一个自包含的.zip文件,整个语料库以425个.tar包的形式分发,总数据量约为800 GB。数据集持续更新,新的DFT计算结果会定期注入,受影响的包会被重新上传。数据集适用于化学、计算化学、DFT、反应机理等领域的研究和应用,特别适合用于开发和分析反应计算模型。每个反应.zip文件内部包含DFT-SinglePoint目录(ORCA单点计算输出,包括.out、.xyz、.inp文件、CHELPG电荷、模糊/Mayer/Wiberg键级矩阵以及status.json)、xTB-scan目录(GFN2-xTB自旋/结构扫描数据,包括每帧几何结构、spin_scan_summary.csv以及bem_snapshots/step_*.json)和IRC_Analysis目录(IRC帧分析及反应摘要),此外还包含轨迹文件(.trj、.xyz)和元数据文件(.yaml)。数据处理过程包括无损清理原始输出,去除冗余信息,但保留了所有科学量(几何结构、能量、电荷、键级、轨道能量)。数据集采用CC-BY-4.0许可证发布。
GoldDIGR is a large-scale open-access DFT reaction calculation dataset containing approximately 1,000,000 xTB/DFT reaction calculations. The data is sourced from the supporting information of published chemistry papers, and recalculated via a unified processing workflow including geometric structure/transition state search, IRC, single-point DFT calculations, and xTB spin/structure scanning. Each calculation is a self-contained .zip file, and the entire corpus is distributed as 425 .tar packages with a total data volume of approximately 800 GB. The dataset is continuously updated, with new DFT calculation results added regularly, and affected packages will be re-uploaded. The dataset is applicable to research and applications in fields such as chemistry, computational chemistry, DFT, and reaction mechanisms, and is particularly suitable for developing and analyzing reaction calculation models. Each reaction .zip file internally contains the DFT-SinglePoint directory (ORCA single-point calculation outputs, including .out, .xyz, .inp files, CHELPG charges, fuzzy/Mayer/Wiberg bond order matrices, and status.json), the xTB-scan directory (GFN2-xTB spin/structure scan data, including per-frame geometric structures, spin_scan_summary.csv, and bem_snapshots/step_*.json), and the IRC_Analysis directory (IRC frame analysis and reaction summaries). Additionally, it includes trajectory files (.trj, .xyz) and metadata files (.yaml). The data processing process includes lossless cleaning of raw outputs to remove redundant information while retaining all scientific quantities: geometric structures, energies, charges, bond orders, and orbital energies. The dataset is released under the CC-BY-4.0 license.
数据集名称
- GoldDIGR (DFT Reaction Calculation Dataset)
规模与格式
- 包含约 1,000,000 个 xTB/DFT 反应计算数据。
- 总大小约 800 GB,分为 425 个
.tar压缩包(每个包含 ≤2500 个.zip文件)。 - 每个计算任务为一个独立的
.zip文件。
数据结构与布局
- 目录结构:
bundles/:存放 425 个.tar包及校验文件(SHA256SUMS、每个包的.sha256校验和)。bundle_manifest.tsv:映射每个.zip到其所属的 bundle。bundle_lists/bundle_XXXX.paths:每个 bundle 内的.zip列表。
.zip文件存储路径:<DOI_prefix>/<DOI_suffix>/<SI_stem>/<reaction>_<charge>_<multiplicity>.zip(例如10.1039/C8SC02758G/c8sc02758g2/03_1_1.zip)。
每个反应 .zip 文件内容
- DFT-SinglePoint/:ORCA 单点计算输出(
.out,.xyz,.inp, CHELPG 电荷、模糊/Mayer/Wiberg 键级矩阵),以及status.json(记录各驻点计算状态)。 - xTB-scan/:GFN2-xTB 自旋/结构扫描数据(每帧几何、
spin_scan_summary.csv各自旋能量、bem_snapshots/每步能量+电荷+键级 JSON)。 - IRC_Analysis/:IRC 轨迹分析(JSON + CSV)及反应总结。
- 其他:TS-opt/IRC 轨迹文件(
.trj,.xyz,.yaml)及元数据。
数据来源与处理
- 从已发表化学论文的支持信息中挖掘原始输出,使用统一流程重新计算(几何优化/过渡态搜索、IRC、单点 DFT、xTB 自旋/结构扫描)。
- 处理过程:无损清理(去除横幅、引用、压缩冗余文件),保留所有科学量(几何、能量、电荷、键级、轨道能量)。
更新机制
- 持续更新,新 DFT 结果增量加入。
- 仅重建和重新上传受影响的 bundle(通过
bundle_manifest.tsv追踪)。
许可证
- CC-BY-4.0。
标签
- chemistry, computational-chemistry, dft, reaction-mechanisms, orca, xtb。




