遇见数据集

Benchmark dataset and reproducibility artefacts for: A symbolic-regression instrument for spectral-exponent estimation in turbulent and critical scale-invariant systems

收藏
Zenodo2026-04-26 更新2026-05-26 收录
官方服务:

资源简介:

Companion dataset to the manuscript "A symbolic-regression instrument for spectral-exponent estimation in turbulent and critical scale-invariant systems: design, characterisation, and uncertainty budget" submitted to SciPost Physics Core (April 2026, single-author: Igor Merlini, ActarusLab). This deposit contains the complete reproducibility package for the symbolic-regression spectral-exponent instrument described in the companion manuscript. It includes all primary and derived numerical results referenced in the paper, plus high-resolution figure files. The dataset comprises five distinct experimental campaigns: (1) Synthetic Rogallo benchmark (Section 3.1): n = 735 configurations spanning p ∈ [0.30, 3.00] at three grid resolutions (N ∈ {64, 96, 128}) and five seeds. Aggregate metrics: MAE = 0.0144, mean bias = +0.0066, R² = 0.9993. (2) Additive-noise stress test (Section 3.2): n = 110 cases across 11 SNR levels from −3 dB to +30 dB plus the noise-free limit, with two reference target exponents (p = 5/3 and p = 0.91). (3) Three-class asymmetry verification (Section 3.3): n = 90 cases verifying the falsifiable physical prediction that recovery error scales monotonically with spectral steepness; confirmed at 4 out of 5 SNR levels. (4) Bootstrap uncertainty quantification (Section 3.4): B = 2000 resamples on the per-seed estimates of the class-separation observable Δ. Clean-limit estimate: Δ = 0.7547, 95% CI = [0.7445, 0.7652]. (5) Direct numerical simulation validation (Section 3.5): in-house pseudo-spectral DNS at N = 96^3, Re_box ≈ 2000, with feedback-controlled stochastic forcing. Recovered exponent on the early steady-state subset: p_DNS = 1.6163 ± 0.0436, consistent with Kolmogorov 1941 (5/3) within the synthetic-benchmark systematic bias. All results are reproducible bit-exactly from the included data tables on dual NVIDIA T4 GPUs in approximately 50 minutes total runtime. Software environment: Python 3.12, JAX 0.4 (CUDA 12), PySR with SymbolicRegression.jl 1.11, Julia 1.11.5. The package is structured in two subdirectories: data/ (CSV + JSON tables) and figures/ (PNG, 200 DPI). Complete schema documentation for every column of every data table is provided in README.md. Released under CC BY 4.0.

本数据集为投稿至《SciPost Physics Core》(2026年4月,单作者:Igor Merlini,ActarusLab)的手稿《用于湍流与临界尺度不变系统谱指数估计的符号回归工具:设计、表征与不确定性预算》的配套数据集。 本存档包包含配套手稿中所述符号回归谱估计工具的完整可复现套件,涵盖论文中引用的全部原始与衍生数值结果,以及高分辨率图像文件。本数据集包含5组独立实验测试: (1) 合成罗加洛基准测试(3.1节):共735组配置,参数p取值范围为[0.30, 3.00],覆盖3种网格分辨率(N∈{64,96,128})与5组随机种子。综合评价指标为:平均绝对误差(MAE)=0.0144,平均偏差=+0.0066,决定系数(R²)=0.9993。 (2) 加性噪声压力测试(3.2节):共110个测试案例,覆盖从-3 dB到30 dB的11种信噪比(SNR)水平,外加无噪声极限场景,包含两组参考目标谱指数(p=5/3与p=0.91)。 (3) 三类不对称性验证(3.3节):共90个测试案例,用于验证“恢复误差随谱陡度单调变化”这一可证伪物理预言;该结论在5种信噪比水平中的4种上得到验证。 (4) 自助法(Bootstrap)不确定性量化(3.4节):针对类分离可观测量Δ的每种子估计执行B=2000次重采样。无噪声极限下的估计值为Δ=0.7547,95%置信区间(CI)为[0.7445, 0.7652]。 (5) 直接数值模拟(DNS)验证(3.5节):采用自研伪谱直接数值模拟方法,网格尺度为N=96³,盒雷诺数Re_box≈2000,采用反馈控制随机强迫。针对早期稳态子集得到的恢复谱指数为p_DNS=1.6163±0.0436,与科尔莫戈罗夫1941理论(5/3)的偏差处于合成基准测试的系统误差范围内。 所有结果均可通过附带的数据表在双NVIDIA T4图形处理器上逐位精确复现,总运行时长约50分钟。软件环境配置为:Python 3.12、JAX 0.4(CUDA 12)、搭载SymbolicRegression.jl 1.11的PySR、Julia 1.11.5。 本套件包含两个子目录:data/(存储CSV与JSON格式数据表)以及figures/(存储200 DPI的PNG格式图像)。所有数据表的每一列均有完整的架构说明文档,收录于README.md文件中。 本数据集采用CC BY 4.0许可协议发布。

提供机构:
Zenodo
创建时间:
2026-04-26
二维码
社区交流群
二维码
科研交流群
商业服务