Gradient Guided Furthest Point Sampling MD17 Data
收藏资源简介:
This dataset contains 630 pickled output files generated from Gradient Guided Furthest Point Sampling (GGFPS) experiments using FCHL representations on seven MD17 molecules (aspirin, benzene, malonaldehyde, napthalene, paracetamol, toluene, uracil). File names encode the molecule, labeled set size (lss), bootstrap replicate index (ind), and target training-set-size group (tss). Each pickle stores a pandas DataFrame with model-evaluation outputs, including per-sample test errors, MAE, selected train/test indices, selected hyperparameters, and timing breakdowns. Training set sizes of 50, 100, 250, and 500 are contained within the tss500 files. The accompanying CSV manifest provides one row per pickle file with metadata, content summary fields, and SHA-256 checksums for integrity verification.



