遇见数据集

Frozen build, harness and locked confirmatory results for the TAALAF Study 2 PI tuning evaluation under identification uncertainty

收藏
Zenodo2026-08-13 更新2026-08-20 收录
官方服务:

资源简介:

Companion deposit to a preregistered simulation study of PI tuning for integrating separator-level loops under identification uncertainty. It contains the registered software bundle, the locked confirmatory outputs, the registered secondary analyses, and the execution log. This deposit follows the data and makes no timing claim. The timing evidence is the OSF registration 10.17605/OSF.IO/S782J, registered 2026-08-10, public with no embargo, which preceded the existence of any confirmatory data. That precedence is not a promise: the confirmatory master seed is a deterministic function of the registration DOI itself, so the dataset could not have been generated before the registration was issued. It must not be confused with the Study 1 pre-registration analysis-code deposit at 10.5281/zenodo.21784512, whose entire value is its timing. Any reader can regenerate the dataset from two public numbers. With basis = 10.17605/osf.io/s782j|fb70ab9d7508c1fb39454bacdac789ad56e7f8ce1a223bff58a3bd40cd647ca1 the digest is c6577af030474c56f8b0424110f600aa8c07138f8978f571d67992df9f31377a and the confirmatory master seed is 1180138225. The first input is the registration DOI, bare and lowercased; the second is the SHA-256 of TAALAF_Study2_Registration_Bundle_v0.19.1.zip, included here and recorded in a visible field of the registration. The runner recomputes the seed and refuses to proceed on mismatch; the run recorded seed_recomputed_and_matched true. Design. 395 scenarios, exactly 79 in each of five disturbance strata. Every policy is applied to every scenario, so all comparisons are paired within scenario. The complete policy set was retuned and re-evaluated independently at M_ST = 1.6 and M_ST = 2.0, both co-reported, with M_ST = 1.4 as the registered sensitivity. A comparative claim is made only where a contrast survives Holm adjustment at both co-reported operating points. Headline results. The abstention guard prevents far more harm than it introduces: at M_ST = 1.6 the SIMC pair shows 62 harms prevented against 3 introduced on 120 activated cases, net value 0.492 with 95% bootstrap interval 0.391 to 0.586, and the OptimalPI pair 56 against 2 on 101 activated, net 0.535 with interval 0.429 to 0.636. p_introduce lies between 0.020 and 0.026 at every operating point. Five of the six headline contrasts are claimed. On deployed harm against SIMC-Guarded the discordant counts are 0 versus 16 at 1.6 and 0 versus 18 at 2.0: AutoTune-Selective-MST was not harmful in a single scenario where SIMC-Guarded was not. One contrast is not claimed. AutoTune-Selective-MST versus OptimalPI-Guarded on deployed harm is not significant at 1.6 and is significant at 2.0 in the direction favouring OptimalPI-Guarded. It is reported as operating-point dependent, exactly as the registered expectation E5 pre-committed, and no claim is made from it in either direction. The oracle's apparent harm is entirely an incumbent artefact. The registered incumbent-comparability sensitivity restricts every harm summary to scenarios whose incumbent itself satisfies the true constraints on hidden truth. Oracle-OptimalPI harm falls from 48/395 to 0/174 at M_ST = 1.6, and from 43/395 to 0/236 at 2.0. Expectation E1 recorded in advance that a non-zero oracle harm rate would be correct behaviour driven by infeasible sluggish incumbents, and would be an implementation verification of the optimizer rather than an empirical finding. It is confirmed exactly. Objective ordering reverses. At 1.6, Oracle-SIMC has mean total valve variation 1052 against Oracle-OptimalPI at 120, but mean IAE 29,280 against 88,996. The arm that minimises valve travel pays roughly three times the regulation error. Expectation E5 pre-committed to reporting exactly this divergence. The validity classifier separates good estimates from bad ones. Median absolute dead-time error is 10 s for accepted estimates against 30 s for rejected, and median relative valve-effect error 4.2 percent against 25.1 percent. The retrospective rejection-error association is AUC 0.70 for dead time and 0.78 for valve effect; this is descriptive, not an operational ROC AUC, because true error is unavailable at deployment time. Abstention costs coverage. Guarded arms issue a controller in 68.9 percent of scenarios against 93.9 to 99.2 percent for unguarded arms. Identification uncertainty also costs realised robustness: at a 1.6 design target the constraint is applied to the estimated model, and realised M_ST on hidden truth exceeds the target in 108 of 395 scenarios for AutoTune-Selective-MST, 136 for SIMC-Guarded and 103 for OptimalPI-Guarded, while the oracle arms, which see the truth, never exceed it. Precision shortfall, recorded in advance. Expectation E6 recorded before generation that guard activation might fall below the 50 percent planning fraction. It was 120/395 and 101/395, that is 25 to 30 percent, and realised guard-value half-widths are approximately 0.10 against the planned 0.0698. Joint coverage at 68.9 percent was above its floor. This is reported as a limitation and is not used to justify a different sample size. Deviations disclosed. First, the registered seed helper writes the lock file as UTF-8 with a byte-order mark, which Node's JSON parser rejects; this was caught before the run, the mark was stripped, no value was altered, and the seed derivation is unaffected. Second, the runner emits the six headline contrasts and the guard-value estimands but not several analyses registered as secondary; those were computed afterwards from the locked runs.json using the registered definitions with no threshold, filter, label or definition altered, and are provided as registered_secondary_analysis.json. Nothing else deviated. The output manifest digest, recorded before any result was opened, is 13c8d415ee61b7dc2c90d89192474537c5f832de17a4de01543972c50bd5bccb. The execution log records the order of events, including the point at which the outputs were hashed and the statement that no result had then been read. The evaluated software is developed and owned by TAALAF Engineering Solutions, the study was funded by that company, and the author holds a commercial interest in it. The funder, the owner of the evaluated tool and the author are not independent parties. The one contrast that could not be claimed under the registered rule is reported as unclaimed even though its direction at the practice-representative operating point is unfavourable to the author's own product.

提供机构:
Zenodo
创建时间:
2026-08-13
二维码
社区交流群
二维码
科研交流群
商业服务