DAN Two Layer Retrievals - Sols 2301-3100
收藏资源简介:
DAN Two-Layer Retrieval Data Release (Sols 2301–3100) This repository contains processed Dynamic Albedo of Neutrons (DAN) two-layer retrieval products and supporting summary products for Curiosity observations spanning sols 2301–3100. The release is organized into region-level summary products and per-location retrieval products. File names use a consistent `site / drive / start_sol / stop_sol` convention so that products from different subdirectories can be matched directly. Retrievals are performed on coadds of all observations at a given site/drive, producing one retrieval per location. Repository structure 2_layer_DAN_sols_2301_3100/├── regional_summaries/└── retrieval_products/ regional_summaries/ This directory contains region-scale summary tables and regional overlay figures. Each region contained in the data volume contains three files. One file stores per-observation summary statistics, the second stores bulk region summary statistics and the third is a plot of the regions retrieval results. Example: - region_<region>_per_observation_stats_snr5_sub50_hdi68.csv- region_<region>_region_summary_stats_snr5_sub50_hdi68.csv- region_<region>_overlay_COMBINED_snr5_sub50.png The same naming pattern is used across regions. retrieval_products/ This directory contains the observation-level retrieval products and diagnostics. retrieval_products/├── background_subtracted_coadd/├── coadd_observations/├── coadd_product/├── corner_plots/├── GMMs/│ ├── unmix_2/│ │ └── gmm_mix/│ └── unmix_variable/│ ├── gmm_mix/│ └── gmm_selection/├── MCMC_backend/├── SNR/├── times/└── walker_plots/ Naming convention Most files share a common stem: site_<SITE>_drive_<DRIVE>_start_sol_<START>_stop_sol_<STOP> For example: site_026_drive_1274_start_sol_542_stop_sol_542This stem is followed by a product-specific suffix, for example: - _bg_dat.npy- _label_matched.txt- _coadded.npy- _corner_plot.png- _gmm_mix.png- _gmm_selection.png- _MCMC.h5.zip- _CETN_SNR.npy- _CTN_SNR.npy- _times.npy- _walker_plot.png This convention allows all products associated with a given coadded retrieval to be aligned by filename. Retrieval content The products in this release correspond to a two-layer retrieval framework. The primary retrieved physical parameters are: - Bottom Layer WEH: bottom-layer water-equivalent hydrogen, in wt%- Top Layer WEH: top-layer water-equivalent hydrogen, in wt%- Bottom Layer Σ_abs: bottom-layer bulk macroscopic neutron absorption cross section (BNACS), in cm^2/g- Top Layer Σ_abs: top-layer bulk macroscopic neutron absorption cross section (BNACS), in cm^2/g- Depth: modeled interface depth between the two layers, in cm Posterior diagnostics may also include an additional logf[counts] fit parameter in the MCMC products. Observation-level products Each retrieval generally includes one file in each of the subdirectories below. background_subtracted_coadd/ Files ending in _bg_dat.npy contain the background-subtracted count data used in retrieval processing. coadd_observations/ Files ending in _label_matched.txt list the DAN observation label or labels contributing to the coadded retrieval product. These files provide the traceability link between the retrieval and the original contributing observation set. coadd_product/ Files ending in _coadded.npy contain the coadded observation-space product used as the retrieval input. corner_plots/ Files ending in _corner_plot.png are posterior diagnostic figures showing parameter distributions and pairwise covariances. GMMs/ This directory contains Gaussian mixture model post-processing products. - unmix_2/ contains fixed two-component mixture summaries.- unmix_variable/ contains variable-component mixture exploration products. Within these directories: - gmm_mix/ contains posterior unmixing summary plots.- gmm_selection/ contains model-selection plots used to compare mixture counts. These products support interpretation of multimodal posterior structure. MCMC_backend/ Files ending in _MCMC.h5.zip are compressed HDF5 backends containing the archived MCMC chains. These are the primary reproducibility products for users who want to regenerate posterior summaries, diagnostics, or alternate post-processing results. SNR/ This directory contains per-observation signal-to-noise products: - *_CETN_SNR.npy- *_CTN_SNR.npy These arrays store time-bin-level SNR values associated with the retrieval input data. times/ Files ending in _times.npy contain the time-bin definitions associated with the DAN die-away measurement. walker_plots/ Files ending in _walker_plot.png show the evolution of the MCMC walkers as a function of iteration number and are intended for quality control and convergence assessment. Regional summary products Per-observation regional statistics Files named region_<region>_per_observation_stats_snr5_sub50_hdi68.csv contain one row per parameter per observation for all observations assigned to a region. These tables provide per-observation posterior summary statistics and identifying metadata such as sol, site, drive, start/stop sol, and source retrieval filename. Coadded observations are only included in regional summaries if their SNR is greater than or equal to 5. Summary statistics include quantities such as: - number of posterior samples- mean- standard deviation- median- KDE mode- percentile-based intervals- mode-centered interval terms- 68% highest-density interval bounds and width These files are intended for downstream analysis, filtering, and custom regional comparisons. Region-level summary statistics Files named region_<region>_region_summary_stats_snr5_sub50_hdi68.csv contain region-level aggregate summaries for each retrieval parameter. These include statistics derived from the distribution of per-observation medians, modes, interval widths, and mixture-distribution summaries. These products are designed for compact regional characterization without requiring users to reprocess all observation-level posterior samples. Regional overlay figures Files named region_<region>_overlay_COMBINED_snr5_sub50.png show the combined regional posterior-density overlays for the two-layer retrieval parameters. These figures provide a visual summary of how posterior structure varies across all observations within a region. Processing tags in filenames Several filenames include processing tags that encode how the products were generated: - snr5: products generated using an SNR threshold of 5- sub50: products generated using a 50% posterior subsampling step- hdi68: summary statistics reported using 68% highest-density intervals These tags are part of the product identity and should be preserved when referencing or redistributing derivative products. Recommended use A typical use pattern is: 1. Start with regional_summaries/ to review region-scale behavior and identify observations of interest.2. Use coadd_observations/ to trace a retrieval back to the contributing DAN observation labels.3. Use coadd_product/, background_subtracted_coadd/, SNR/, and times/ for observation-space analysis.4. Use corner_plots/, walker_plots/, and GMMs/ for posterior interpretation and quality control.5. Use MCMC_backend/ when full reproducibility or custom post-processing is required. Acknowledgment If you use these products in published work, please cite the associated data release and the relevant scientific publications describing the DAN retrieval methodology.
DAN双层反演数据集发布(Sol 2301–3100) 本仓库包含经预处理的中子动态反照率(Dynamic Albedo of Neutrons, DAN)双层反演产物及配套汇总产物,覆盖好奇号火星车Sol(火星日)2301至3100期间的观测数据。本次发布按区域级汇总产物与单点位反演产物两类组织。文件名采用统一的`site / drive / start_sol / stop_sol`命名规范,可直接匹配不同子目录下的关联产物。反演基于给定站点/行进路线下所有观测的叠加数据完成,每个点位生成一份独立反演结果。 ### 仓库结构 2_layer_DAN_sols_2301_3100/ ├── regional_summaries/ └── retrieval_products/ #### regional_summaries/ 该目录包含区域级汇总表格与区域叠加可视化图。数据卷中的每个区域均包含三类文件:第一类存储单观测汇总统计量,第二类存储区域整体汇总统计量,第三类为该区域反演结果的可视化图。示例如下: - region_<region>_per_observation_stats_snr5_sub50_hdi68.csv - region_<region>_region_summary_stats_snr5_sub50_hdi68.csv - region_<region>_overlay_COMBINED_snr5_sub50.png 所有区域均采用统一命名模式。 #### retrieval_products/ 该目录包含观测级反演产物与诊断数据,其子目录结构如下: retrieval_products/ ├── background_subtracted_coadd/ ├── coadd_observations/ ├── coadd_product/ ├── corner_plots/ ├── GMMs/ │ ├── unmix_2/ │ │ └── gmm_mix/ │ └── unmix_variable/ │ ├── gmm_mix/ │ └── gmm_selection/ ├── MCMC_backend/ ├── SNR/ ├── times/ └── walker_plots/ ### 命名规范 多数文件共享通用前缀: `site_<SITE>_drive_<DRIVE>_start_sol_<START>_stop_sol_<STOP>` 示例:`site_026_drive_1274_start_sol_542_stop_sol_542` 该前缀后接产物专属后缀,例如: - _bg_dat.npy - _label_matched.txt - _coadded.npy - _corner_plot.png - _gmm_mix.png - _gmm_selection.png - _MCMC.h5.zip - _CETN_SNR.npy - _CTN_SNR.npy - _times.npy - _walker_plot.png 该命名规范可通过文件名直接匹配对应叠加反演的所有关联产物。 ### 反演内容 本次发布的产物对应双层反演框架,主要反演的物理参数包括: - 下层水等效氢(Bottom Layer WEH):下层水等效氢含量,单位为重量百分比(wt%) - 上层水等效氢(Top Layer WEH):上层水等效氢含量,单位为重量百分比(wt%) - 下层宏观中子吸收截面总和(Bottom Layer Σ_abs):下层整体宏观中子吸收截面(bulk macroscopic neutron absorption cross section, BNACS),单位为cm²/g - 上层宏观中子吸收截面总和(Top Layer Σ_abs):上层整体宏观中子吸收截面(BNACS),单位为cm²/g - 层厚(Depth):两层界面的模拟深度,单位为cm 后验诊断产物中还可能包含马尔可夫链蒙特卡洛(Markov Chain Monte Carlo, MCMC)产物内的额外`logf[counts]`拟合参数。 ### 观测级产物 每个反演通常在以下各子目录中对应一份文件: 1. **background_subtracted_coadd/**:以`_bg_dat.npy`结尾的文件包含反演处理中使用的背景扣除计数数据。 2. **coadd_observations/**:以`_label_matched.txt`结尾的文件列出了参与该叠加反演的DAN观测标签,可建立反演与原始贡献观测集之间的可追溯关联。 3. **coadd_product/**:以`_coadded.npy`结尾的文件包含用作反演输入的观测空间叠加产物。 4. **corner_plots/**:以`_corner_plot.png`结尾的文件为后验诊断图,展示参数分布与两两协方差。 5. **GMMs/**:该目录包含高斯混合模型(Gaussian Mixture Model, GMM)后处理产物: - `unmix_2/`:固定双组分混合模型的汇总结果 - `unmix_variable/`:可变组分混合模型的探索产物 在这些子目录中: - `gmm_mix/`:包含后验解混汇总图 - `gmm_selection/`:包含用于比较混合组分数量的模型选择图 此类产物可辅助解释多峰后验结构。 6. **MCMC_backend/**:以`_MCMC.h5.zip`结尾的文件为压缩HDF5格式的后端文件,包含存档的MCMC链,是希望复现后验汇总、诊断或自定义后处理结果的用户的核心可复现产物。 7. **SNR/**:该目录包含单观测信噪比(Signal-to-Noise Ratio, SNR)产物: - `*_CETN_SNR.npy` - `*_CTN_SNR.npy` 这些数组存储与反演输入数据相关的时间分箱级信噪比数值。 8. **times/**:以`_times.npy`结尾的文件包含与DAN衰减测量相关的时间分箱定义。 9. **walker_plots/**:以`_walker_plot.png`结尾的文件展示MCMC采样器随迭代次数的演化过程,用于质量控制与收敛性评估。 ### 区域汇总产物 #### 单观测区域统计量 以`region_<region>_per_observation_stats_snr5_sub50_hdi68.csv`命名的文件中,每一行对应一个观测的每个参数,涵盖所有分配至该区域的观测。此类表格提供单观测后验汇总统计量与标识元数据,包括Sol、站点、行进路线、起始/终止Sol以及源反演文件名。仅当叠加观测的信噪比不低于5时,才会被纳入区域汇总。 汇总统计量包含以下指标: - 后验样本数量 - 均值 - 标准差 - 中位数 - 核密度估计(Kernel Density Estimation, KDE)众数 - 基于百分位数的区间 - 众数中心区间项 - 68%最高密度区间(Highest Density Interval, HDI)的边界与宽度 此类文件可用于下游分析、筛选与自定义区域对比。 #### 区域级汇总统计量 以`region_<region>_region_summary_stats_snr5_sub50_hdi68.csv`命名的文件包含每个反演参数的区域级聚合汇总结果,涵盖基于单观测中位数、众数、区间宽度与混合分布汇总的统计量。 此类产物旨在实现紧凑的区域特征描述,无需用户重新处理所有观测级后验样本。 #### 区域叠加图 以`region_<region>_overlay_COMBINED_snr5_sub50.png`命名的文件展示了双层反演参数的组合区域后验密度叠加图,可可视化区域内所有观测的后验结构差异。 ### 文件名中的处理标签 部分文件名包含处理标签,用于编码产物的生成方式: - `snr5`:采用信噪比阈值5生成的产物 - `sub50`:采用50%后验抽样步骤生成的产物 - `hdi68`:采用68%最高密度区间报告的汇总统计量 此类标签属于产物标识的一部分,在引用或分发衍生产物时应予以保留。 ### 推荐使用流程 典型的使用模式如下: 1. 从`regional_summaries/`目录入手,查看区域尺度的行为特征并识别感兴趣的观测 2. 使用`coadd_observations/`目录将反演追溯至对应的贡献DAN观测标签 3. 使用`coadd_product/`、`background_subtracted_coadd/`、`SNR/`与`times/`目录开展观测空间分析 4. 使用`corner_plots/`、`walker_plots/`与`GMMs/`目录开展后验解释与质量控制 5. 当需要完全复现或自定义后处理时,使用`MCMC_backend/`目录的产物 ### 致谢说明 若您在已发表工作中使用此类产物,请引用相关数据集发布文章与描述DAN反演方法的相关科学出版物。



