Fractal Aggregate Collision-Efficiency Dataset, Graph Representations, Model Checkpoints, and Analysis Outputs for Dual-Scale Graph Learning
收藏资源简介:
This record contains the complete data, trained models, analysis outputs, and source code supporting the study “Global morphology and particle-resolved organization jointly predict random-walk collision efficiency of fractal aggregates.” The dataset comprises 50,000 synthetic rigid three-dimensional fractal aggregates containing 50–500 primary particles. Each aggregate includes particle coordinates, contact-graph connectivity, particle count, and a finite-trajectory Monte Carlo collision-efficiency label calculated using 2,000 off-lattice random walkers with a maximum trajectory length of 25,000 steps. The deposit includes the raw aggregate batches, processed graph representations, five seed-specific train–validation–test split manifests, feature-standardization records, trained checkpoints for all benchmark and representation-isolation models, training histories, held-out predictions, and numerical results underlying the manuscript and Supplementary Information. The deposited analyses cover model benchmarking, representation isolation, Monte Carlo repeatability, residual and morphology diagnostics, rigid-transformation invariance, repeated descriptor permutation, attention-based deletion, graph-centrality comparisons, generator-quality assessment, overlap-impact testing, trajectory-cap sensitivity, finite-horizon robustness, and computational profiling. The source-code archive contains the scripts used for aggregate generation, graph preprocessing, model training, prediction, statistical evaluation, interpretability analysis, geometric auditing, trajectory continuation, and runtime profiling, together with the original Google Colab generator and supporting provenance records. A strict geometric audit identified 49,991 quality-passing aggregates and nine structures containing center-to-center overlaps below the specified threshold; the affected identifiers and corresponding exclusion analyses are included. The deposited raw batch files constitute the authoritative dataset record.



