遇见数据集

Benchmark data and scripts: MACE-MP-0 evaluation cost, per-task orchestration overhead, and GPU energy for 120 hypothetical metal-organic frameworks

收藏
Zenodo2026-09-27 更新2026-10-01 收录
官方服务:

资源简介:

This dataset accompanies the manuscript "Infrastructure Challenges in Scaling AI-Driven Materials Discovery: A Critical Review" by N. Goel, T. Das, and W. A. Goddard III (submitted to Computational Materials Science). It contains the structures, scripts, and raw data for a controlled single-node measurement comparing machine-learning interatomic potential (MLIP) evaluation cost with per-task workflow overheads, and the GPU energy of MLIP relaxations, for hypothetical metal-organic frameworks (MOFs). CONTENTSmofs/ - 120 hypothetical MOFs (CIF) generated with PORMAKE 0.2.3 (random seed 0), 30 in each of four size bins (0-300, 300-600, 600-1000, 1000-2000 atoms per cell), spanning 99 topologies. Candidates were restricted to single-node, single-edge topologies with metal-containing nodes fitting the topology (RMSD < 0.3 Å) and organic linkers. Structures with any interatomic distance below 0.8 Å were discarded (120 kept from 5,627 candidates). manifest.csv lists each structure; generation_stats.txt records the generation statistics. scripts/ - generate_mofs.py (structure generation); bench_mace.py (per-step timing and FIRE relaxations); bench_parsl.py (Parsl dispatch latency and throughput); bench_startup.py (cold-start cost of a fresh Python process); bench_energy.py (GPU energy and utilization); analyze.py and analyze_energy.py (tables and figures); env_freeze.txt (exact Python environment). results/ - mace_results.csv (per-structure timings and relaxation outcomes); startup.csv (cold-start timings); parsl_w1.json and parsl_w10.json (Parsl benchmarks with 1 and 10 workers); energy_results.csv (per-structure GPU energy); energy_idle.json (idle power); gpu_util.csv (GPU utilization and power samples); node_info.txt (hardware and driver details). METHODSHardware: one NVIDIA V100-SXM2-32GB GPU, 2x Intel Xeon E5-2698 v4 CPUs (Caltech HPC). Software: Python 3.11.16, PyTorch 2.5.1 (CUDA 12.1 build), mace-torch 0.3.16, ASE 3.29.0, Parsl 2026.09.14, PORMAKE 0.2.3. MLIP: MACE-MP-0 small in float32 through its ASE calculator, loaded once and kept in GPU memory. For each MOF, 20 energy and force evaluations were timed after 3 discarded warm-up evaluations, with GPU synchronization before and after each call, followed by one FIRE relaxation of atomic positions at fixed cell (fmax = 0.05 eV/Å, at most 500 steps). Overheads: Parsl dispatch latency is the median round-trip time of 1,000 sequential no-op tasks with the High-Throughput Executor; throughput is the completion rate of 10,000 concurrently submitted no-op tasks. Cold-start cost was measured in 10 fresh Python processes, each importing PyTorch, initializing CUDA, importing MACE and ASE, loading the model, and evaluating the smallest MOF. Energy: measured in a second run with the same protocol, using the GPU's cumulative energy counter (NVML) read before and after each relaxation; GPU utilization and power were sampled every 0.1 s. Idle power was measured over 60 s with the model loaded. Energy values cover the GPU only. NOTESThe structures are hypothetical and computer-generated, not experimentally known materials. The MACE-MP-0 model weights are not redistributed here. They are downloaded automatically by mace-torch. Usage instructions are in README.md.

提供机构:
Zenodo
创建时间:
2026-09-27
二维码
社区交流群
二维码
科研交流群
商业服务