modelforge curated dataset: SPICE 2 OpenFF
收藏资源简介:
Modelforge Curated SPICE 2 OpenFF Dataset:- Full dataset- Version: full_dataset_v1.1 This provides a curated hdf5 file for the SPICE 2 OpenFF dataset (Open Force Field initiative default level of theory) designed to be compatible with modelforge, an infrastructure to implement and train NNPs. This dataset contains 112628 unique records for 1971769 total configurations. This excludes any configurations where the magnitude of any forces on the atoms are greater than 1 hartree/bohr. When applicable, the units of properties are provided in the datafile, encoded as strings compatible with the openff-units package. This is compatible with modelforge HDF5 schema 2. For more information about the structure of the data file, please see the following: https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module Properties Included: atomic_numbers positions "per_atom" "nanometer" dft_force "per_atom" "kilojoule_per_mole / nanometer" mbis_charges "per_atom" "elementary_charge" dispersion_correction_force "per_atom" "kilojoule_per_mole / nanometer" dft_total_force "per_atom" "kilojoule_per_mole / nanometer" total_charge "per_system" "elementary_charge" dft_energy "per_system" "kilojoule_per_mole" scf_dipole "per_system" "elementary_charge * nanometer" dispersion_correction_energy "per_system" "kilojoule_per_mole" dft_total_energy "per_system" "kilojoule_per_mole" source "meta_data" molecular_formula "meta_data" canonical_isomeric_explicit_hydrogen_mapped_smiles "meta_data" Source Dataset: Small-molecule/Protein Interaction Chemical Energies (SPICE). The SPICE dataset contains 1.1 million conformations for a diverse set of small molecules, dimers, dipeptides, and solvated amino acids. It includes 17 elements (H, Li, B, C, N, O, F, Na, Mg, Si, P, S, Cl, K, Ca, Br, I), charged and uncharged molecules, and a wide range of covalent and non-covalent interactions. It provides both forces and energies calculated using B3LYP-D3BJ/DZVP level of theory, using Psi4 1.4.1.This is the default theory used for force field development by the Open Force Field Initiative. SPICE 2 builds upon SPICE 1 (i.e., adding new collections to the dataset). It includes the following collections from the MolSSI qcarchive: Collections as part of SPICE 1 "SPICE Solvated Amino Acids Single Points Dataset v1.1", "SPICE Dipeptides Single Points Dataset v1.2", "SPICE DES Monomers Single Points Dataset v1.1", "SPICE DES370K Single Points Dataset v1.0", "SPICE PubChem Set 1 Single Points Dataset v1.2", "SPICE PubChem Set 2 Single Points Dataset v1.2", "SPICE PubChem Set 3 Single Points Dataset v1.2", "SPICE PubChem Set 4 Single Points Dataset v1.2", "SPICE PubChem Set 5 Single Points Dataset v1.2", "SPICE PubChem Set 6 Single Points Dataset v1.2", Note, this does not include the following collections which are part of the standard SPICE 1 and 2 datasets: "SPICE Ion Pairs Single Points Dataset v1.1", "SPICE DES370K Single Points Dataset Supplement v1.0", New collections for SPICE 2: 'SPICE PubChem Set 7 Single Points Dataset OpenFF v1.0' 'SPICE PubChem Set 8 Single Points Dataset OpenFF v1.0' 'SPICE PubChem Set 9 Single Points Dataset OpenFF v1.0' 'SPICE PubChem Set 10 Single Points Dataset OpenFF v1.0' 'SPICE Water Clusters OpenFF v1.1' 'SPICE Amino Acid Ligand OpenFF v1.0' 'SPICE Solvated PubChem Set 1 OpenFF v1.0' 'SPICE PubChem Boron Silicon OpenFF v1.0' Citations: Reference to SPICE 2 publication: Eastman, P., Pritchard, B. P., Chodera, J. D., & Markland, T. E Nutmeg and SPICE: models and data for biomolecular machine learning. Journal of chemical theory and computation, 20(19), 8583-8593 (2024). https://doi.org/10.1021/acs.jctc.4c00794 Original SPICE 1 publication: Eastman, P., Behara, P.K., Dotson, D.L. et al. SPICE, A Dataset of Drug-like Molecules and Peptides for Training Machine Learning Potentials. Sci Data 10, 11 (2023). https://doi.org/10.1038/s41597-022-01882-6



