Research code and data for: A score-based particle flow filter for non-Gaussian data assimilation in high-dimensional chaotic systems
收藏资源简介:
Research code and data for: A score-based particle flow filter for non-Gaussian data assimilation in high-dimensional chaotic systems This repository contains the complete research code, trained model weights, and data generation scripts accompanying the manuscript: Shen, Z., Tang, Y., & Fang, Y. (2026). A score-based particle flow filter for non-Gaussian data assimilation in high-dimensional chaotic systems. arXiv:2608.22454 [physics.ao-ph]. https://arxiv.org/abs/2608.22454 Overview The Score-based Particle Flow Filter (Score-PFF) integrates deep neural network-learned score functions with the particle flow filter framework to enable computationally efficient, non-Gaussian data assimilation. This archive provides all materials necessary to reproduce the figures and numerical experiments reported in the paper, conducted on the 40-dimensional and 1000-dimensional Lorenz-96 systems. Contents - figure_scripts: Plotting scripts that generate the six main paper figures from .npz data files.- experiment_scripts: Scripts to rerun all assimilation experiments (linear and nonlinear observations, 40D and 1000D configurations).- training: Model training code, training data generators, and trained score network checkpoints (MLP, FNO, GNN for 40D; GNN for 1000D).- data: Pre-computed .npz data files consumed by the figure scripts.- figures: Final PNG figures as they appear in the manuscript.- checkpoints: Root-level copies of selected score network weights. Key Components Paper Figure 1 (Prior gradient advantage): Data generated on-the-fly. Reproduction script: figure_scripts/generate_figure1_score_prior_advantage.py Paper Figure 2 (Statistical comparison): Data from data/final_statistical_data.npz. Reproduction script: experiment_scripts/run_20_trials_experiment.py Paper Figure 3 (Q-Q evolution): Data from data/qq_evolution_stable.npz. Reproduction script: experiment_scripts/compare_non_gaussianity_3methods.py Paper Figure 4 (Nonlinear observations): Data from data/nonlinear_obs_*.npz. Reproduction script: experiment_scripts/nonlinear_obs_statistical_20trials.py Paper Figure 6 (Architecture comparison): Data from data/architecture_comparison_linear_sigma05.npz. Reproduction script: figure_scripts/compare_architectures_linear_sigma05.py Paper Figure 7 (1000D scalability): Data from data/final_summary_20runs.npz. Reproduction script: experiment_scripts/compare_final_summary.py Datasets All simulation data are deterministically generated from the Lorenz-96 model using fixed random seeds. 40-dimensional system: 1 million state samples (500 trajectories x 2000 steps, after 1000-step spin-up). 1000-dimensional system: 100,000 state samples (50 trajectories x 2000 steps, after 2000-step spin-up). Run training/generate_training_data.py or training/generate_data_1000d.py with documented random seeds to reproduce the exact datasets. Quick Start To reproduce a figure from the provided data: cd pack_for_submissioncp data/final_statistical_data.npz .python figure_scripts/plot_aligned_comparison.py To regenerate experiment data from scratch: cd pack_for_submissionpython experiment_scripts/run_20_trials_experiment.py To retrain score networks: cd pack_for_submission/trainingpython data_generator.pypython train_score.py --data_path l96_training_data.npz --save_dir checkpoints See README.md in the repository for detailed installation instructions and step-by-step reproduction guides. System Requirements - Python >= 3.10- PyTorch >= 2.0 (CPU or CUDA)- NumPy, SciPy, Matplotlib The 1000-dimensional experiments were originally conducted on an Intel Xeon Platinum 8276 processor with 256 GB system memory. GPU acceleration (NVIDIA L40, 48 GB) is supported for score network inference. Notes - All experiments fix numpy/torch random seeds internally; results are deterministic up to BLAS/threading nondeterminism.- Timing results (Figure 7) are hardware-dependent; RMSE results are hardware-independent.- Large intermediate raw files (e.g., experiment_results_20trials.pkl) are not included but can be regenerated with the provided experiment scripts. Citation If you use this code or data, please cite: Shen, Z., Tang, Y., & Fang, Y. (2026). A score-based particle flow filter for non-Gaussian data assimilation in high-dimensional chaotic systems. arXiv:2608.22454 [physics.ao-ph]. https://arxiv.org/abs/2608.22454



