A Harmonized Lunar Boulder Database Across 46 Study Regions: Size Distributions, Geomorphic Context, and Surface-Exposure Metrics
收藏资源简介:
This dataset and reproducibility archive accompanies the manuscript “Cross-Regional Variability of Lunar Boulder Size Distributions and Threshold-Dependent Surface Exposure.” The archive provides a provenance-preserving harmonization of 108,853 lunar boulder records compiled from seven source datasets and 46 study regions. After quality control, removal of 112 records with invalid size information, and conservative removal of 37 confirmed duplicate records, the analytical database contains 108,704 unique valid boulder records. Original source identifiers, study-region assignments, diameter definitions, and available spatial information are retained to support traceability. The archive includes: standardized record-level and conservatively deduplicated boulder databases; a duplicate-audit log and field-level data dictionary; metadata and source provenance for all 46 study regions; region-specific empirical operational lower fitting thresholds; power-law and shifted-exponential maximum-likelihood estimates; nonparametric-bootstrap parameter confidence intervals; AICc-based relative model comparisons; 2,000-replicate parametric-bootstrap Kolmogorov–Smirnov goodness-of-fit tests; machine-readable results underlying Tables S7 and S11–S13; leave-one-source-out sensitivity analyses; independent KS-minimized threshold sensitivity analyses; empirical CDF, fitted-CDF, residual, and P–P diagnostic data; area-normalized boulder-exposure indicators, including exact counts above 5 m and Garwood exact Poisson confidence intervals; figure-source data and exploratory North Ray spatial-exposure products; reproducible Python scripts, configuration files, software requirements, and SHA-256 checksums. Formal regional model comparison is restricted to 26 regions retaining at least 100 boulders at or above their operational threshold. Relative AICc comparison favors the power-law model in 21 regions and the shifted-exponential model in five. Parametric-bootstrap goodness-of-fit tests indicate that at least one candidate model is not rejected in 12 regions, whereas both candidates are rejected in 14 regions. These results distinguish relative model preference from adequate absolute fit.Operational thresholds are empirical distribution-turnover limits rather than independently calibrated detection-completeness thresholds. The source datasets differ in image resolution, illumination, mapping procedure, diameter definition, and spatial coverage. Deduplication is deliberately conservative, and unresolved physical duplicates cannot be excluded where boulder-level geographic information is unavailable. Absolute-density and exposure analyses are limited to regions with defensible complete mapped areas. Spatial products use observed coordinates and should not be interpreted as engineering-certified hazard maps.The harmonized data are released under CC BY 4.0. Analysis code is provided under the MIT License. Users should cite this Zenodo record, the associated article when available, and the relevant original source publications documented in the archive.



