Reproducibility archive for: A Discrete Synthetic Benchmark for PCA-Ridge and Reference-Ray 3-D Velocity Reconstruction from First-Arrival Travel Times
收藏资源简介:
Reproducibility archive for the article “A Discrete Synthetic Benchmark for PCA-Ridge and Reference-Ray 3-D Velocity Reconstruction from First-Arrival Travel Times”. Learned methods for seismic travel-time tomography are usually compared with physics-based inversion on benchmarks whose representation and evaluation choices are left implicit. We show that the apparent workflow ordering depends materially on the declared representation choices, not only on the estimator. The benchmark reconstructs 3-D P-wave velocity from first-arrival travel times: 250 synthetic geological targets in five families, 384 source-receiver observations each, split 175/37/38 by target, with travel times from a Fast-Sweeping calculation on a 1.25-km grid and errors measured against analytic velocity at 2.5-km cell centers. With corpus and endpoint fixed, the declared identifier-ordered principal component analysis ridge (PCA-ridge) workflow gives root-mean-square error 0.353 km/s, geometry-canonical re-ordering with penalty reselection gives 0.310 km/s, and a cell-native parameterization gives 0.261 km/s, against 0.272 km/s for an intentionally non-equivalent Reference-ray baseline (the cell-native paired interval spans zero); a small neural network in place of the linear map gives 0.352 km/s. The main reported comparison, retrospectively designated and descriptive, is a paired difference of +0.0816 km/s (95% percentile-bootstrap interval [+0.0492, +0.120] km/s) favoring the baseline. Benchmarks should therefore declare encoding and parameterization explicitly. Conclusions are conditional on the declared operator and representation, and establish no field-scale accuracy or universal method ranking. What is here records/ — the evidence base. Every value published in the article and its Supplementary Material re-derives from these files, and the records/... identifiers cited in both documents are paths into this directory. reproducibility/ — SHA256SUMS.txt pins every file by checksum; regeneration_instructions.md gives the commands that rebuild the corpus, records and figures; source/ is the complete analysis package and its 8 entry points; dependency_versions.txt and environment_metadata.txt record the environment of the reported run. qa/ — the verification apparatus. qa/audit_scripts/ holds three suites that re-derive every published value from records/, check that both documents state shared contracts identically, and confirm that every quantified claim is registered in qa/claim_register.json and backed by evidence. They are runnable: a reader can confirm the article's numbers without taking them on trust. manuscript/ and supplementary/ — the article and Supplementary Material sources, in markdown. configs/, figures/, supplementary_figures/ — the frozen production configuration and the figure sources. Why the manuscript is deposited as markdown The markdown sources are not drafts. They are the inputs to a build-and-verify chain: the submitted DOCX and PDF are generated from them, and the build is pinned to their SHA-256 checksums, so the published article is provably the deposited source rather than a separate copy of it. They are also what the verification suites read — the claim layer checks sentences against evidence, which requires the sentences. Depositing the rendered PDF alone would break both properties. The archive holds 2639 files (377 MB). reproducibility/SHA256SUMS.txt pins 2638 of them by SHA-256 — every file except the manifest itself, which cannot contain its own checksum — and reproducibility/README.md is the entry point. Changes in version 1.1.0 This version accompanies the revised article. The manuscript, the Supplementary Material and the captions are revised. Figure 1 and Supplementary Figures S1, S4 and S5 were redrawn at print width from unchanged data; qa/figure_hash_audit.md records the old and new checksums. A shallow MLP coefficient-map baseline was added under records/secondary_diagnostics/mlp_coefficient_map/, with its script. The verification suites, the claim register and the prior-work comparison grid in qa/ were extended or corrected, and reproducibility/regeneration_instructions.md covers the new scripts. No corpus, forward-label, split, estimator or baseline record has changed. The records under records/shuffled_time/ belong to a 100-seed shuffled-time sensitivity that was withdrawn in the revision: an indexing error in its script meant it did not reassign the travel-time vectors as described. They are retained unchanged, so that the earlier version's results remain checkable, and the revised article does not cite them. The single-permutation shuffled-time control reported in the article is unaffected. Records in qa/ written before the revision that list the 100-seed sensitivity predate its withdrawal.



