OPAL: complete experiment data and evidence for authorizing adaptive data acquisition
收藏资源简介:
OPAL: complete experiment data and evidence for authorizing adaptive data acquisition Metadata field Value Record title OPAL: complete experiment data and evidence for authorizing adaptive data acquisition Record https://doi.org/10.5281/zenodo.21859443 Version 1.3.0 Resource type Dataset Release date 10 August 2026 Related article Testing when adaptive data acquisition can replace fixed measurement plans Data licence CC BY 4.0 for author-generated data and documents Software licence Pending author or institutional decision before public software release Creators Jia Bi. Scientific Computing Department, Science and Technology Facilities Council, Rutherford Appleton Laboratory, Didcot, UK. ORCID 0000-0002-3773-3289 Samuel Pinilla. Diamond Light Source, Harwell Science and Innovation Campus, Didcot, UK. Chenyang Zhu. School of Electronics and Computer Science, University of Southampton, Southampton, UK. ORCID 0000-0002-2145-0559 Description to paste into Zenodo Across experimental science, an early measurement increasingly determines which samples receive costly follow-up. This choice affects both resource use and the evidence available to later decisions. Learned selection rules can rank promising measurements, but predicted value alone does not establish that a rule is reliable enough to replace a fixed plan. This record accompanies the article ‘Testing when adaptive data acquisition can replace fixed measurement plans’. The study introduces the opportunity-aware protocol for authorizing learned measurement rules (OPAL). OPAL learns a selection rule from labelled data in the intended setting, fixes that rule before final outcomes are opened and evaluates it on held-out samples. The method separates two questions that are often combined: which follow-up measurement appears valuable, and whether the available evidence is sufficient to change the measurement plan. The article establishes two evidence boundaries. First, outcomes from another setting and unlabelled measurements from the intended setting cannot determine authorization under unrestricted outcome shift. Second, an exact finite-population bound identifies when a favourable pilot is too small to support claims about the campaign’s unmeasured remainder. The principal archive-based evaluation uses 11,265 held-out compounds from the public cpg0012 Cell Painting collection. The prespecified expected-net-benefit rule had the largest held-out executed value, selected 96.01% of the library for additional imaging and had a 97.14% false-activation upper bound. OPAL selected 595 compounds (5.28%), corresponding to 1,190 potential additional wells. Its false-activation upper bound was 5.18%, and its executed-value lower bound remained positive after measurement cost. No other rule that selected compounds had a false-activation upper bound below 37.94%. Uncertainty about errors among selected compounds nevertheless remained above its registered target, so the fixed plan was retained. The result distinguishes a promising measurement branch from evidence sufficient to use it. The record also contains the complete author-generated evidence for the finite-campaign analyses, executed-value and information-cost study, held-out simulator validation, CTRP measured-response analysis, Causal Chambers constructibility analysis, post-hoc plate-block sensitivity and fixed-assignment cost sensitivity. It includes result-bearing unit, bank, campaign and fixture ledgers; frozen assignments and configurations; model and threshold records; outcome and evaluation tables; bootstrap and exact-enumeration outputs; figure source data; reconstruction software; manuscript sources; and integrity manifests. Large third-party source archives are not redistributed. Instead, the package records their public locations, versions or object identifiers, checksums where available, licences and the transformations used to create the author-derived evidence. The included simulation studies are author generated. The Cell Painting and CTRP studies are offline analyses of public experimental archives. They are not prospective instrument deployments. Package contents · Current manuscript, Supplementary Information, cover letter, references, all seven main figures, all nine Supplementary figures, complete compact figure source data and deterministic figure builders. · Seven controlled result-bearing evidence archives covering cpg0012, the executed-value and information-cost analysis, finite-campaign evidence, held-out simulator validation and comparators, CTRP measured outcomes and Causal Chambers constructibility. · Complete cpg0012 post-hoc plate-block and fixed-assignment cost-sensitivity inputs, scripts and outputs. · Reconstruction software, configuration files, frozen thresholds and assignments, analysis ledgers, validation reports and package-level checksums. · Third-party source manifests that identify external archives without redistributing restricted or very large raw files. Licence and reuse notes Author-generated data and documents: CC BY 4.0. Third-party materials: the original licence and attribution of each source applies. See the third-party source manifests in the package. Software: the public software licence is pending an explicit author or institutional decision. Do not describe the code as open source until a top-level LICENSE has been approved and added. Suggested keywords adaptive data acquisition; adaptive experimentation; experimental design; statistical decision-making; safe policy learning; finite-population inference; Cell Painting; high-throughput measurement Related resources Article code: GitHub source repository Evidence record: Zenodo record 21859443



