omicau benchmark results: retained outputs, deviations and reproducibility record for an evaluation with a prospectively archived core design
收藏资源简介:
This record contains the retained outputs of a benchmark of omicau, a multi-omics audit and fusion workflow. The core evaluation design, hypotheses, estimands, thresholds, decision rules and deterministic seed-derivation specification were archived before definitive execution at https://doi.org/10.5281/zenodo.21779751. That protocol package is timestamped and member-manifest-valid, but it was not a complete clean-tree reproducibility freeze: it omitted the expanded seed registry and overlap audit, protocol checksums, environment captures and split manifests required by its own expected-output inventory; its recorded source commit did not contain the protocol directory; and its aggregate content hash is not reproducible. DEV-010 documents these limitations. Public commit 04da2f1a328ea7a9adbd8b67c4e77a4e2a81f64e retained the full seed registry and audit, checksums, generator and harness code, and an environment capture before the first definitive simulation fit. The campaign completed 3,710 independent simulated datasets across seven registered families (38,450 result rows, 0 simulation-method failures), a paired analysis of one external comparator on a retained 240-dataset subset with 4,000 dataset-bootstrap draws, and one real cohort of 738 subjects with matched RNA-seq and DNA methylation evaluated under repeated group-aware cross-validation. Included are retained split indices and hashes, per-dataset result shards, failure records, derived digests, eligibility assessments, acquisition and harmonization provenance, figures, tables, and the code used to generate them. 15 deviation records disclose every identified departure or retrospective limitation. DEV-009 records that the registered C9/H6 resource endpoint is incomplete: timings were observed during the prespecified concurrent six-worker schedule, no isolated timing pass or complete resources.parquet was retained, and external resource comparison is not evaluable. DEV-011 through DEV-015 document the corrected predictive multiplicity families, the non-executable H2 rate-multiplicity clause, the conflicting role-recovery estimands, incomplete real-cohort secondary outcomes and the conflicting real-cohort inferential specifications. After results were available, a post-campaign execution-integrity checker regenerated all 3,710 simulated datasets using the pre-run seed registry, committed generator implementation, protocol inputs and recorded environment; all 38,450 value-level provenance hashes matched. Of 34,740 predictive metrics, 34,690 were exact and 27 additional non-identical rows were within the post-campaign, non-registered 10⁻⁹ checker threshold. Of the 23 above-threshold rows, seven were within post-hoc retained-float32 equal-score tie upper bounds and 16 exceeded those bounds and remain unresolved by that diagnostic; neither group is an H7 tolerance pass. The registered H7 all-applicable-checks rule was not satisfied because the provenance-keyed cache test had no retained evidence. No raw cohort matrix, out-of-fold prediction vector or direct TCGA barcode is included. Public GDC file identifiers are retained for acquisition provenance, but the deposit removes the per-file participant field, aliquot field and pseudonymized local filename that could otherwise link those public identifiers to retained subject pseudonyms. Retained split manifests use salted subject pseudonyms, and the salt is not archived.



