D2A-HWY-60: A Reproducible Benchmark for Uncertainty-Aware Design-to-As-Built Verification in Highway Construction
收藏资源简介:
D2A-HWY-60 is a reproducibility package and controlled benchmark for uncertainty-aware design-to-as-built verification in highway construction. It accompanies the manuscript “An Uncertainty-Aware Digital-Twin Framework for Automated Design-to-As-Built Verification in Highway Construction.” The benchmark contains 60 research-constructed highway scenarios spanning horizontal alignment, vertical profile/elevation, lane and shoulder width, cross-slope/superelevation, surface deviation, and compound nonconformities. The package includes machine-readable LandXML design corridors, structured synthetic as-built evidence, encoded QA/QC rule profiles, deterministic verification outputs, specification-sensitivity results, and station-addressable evidence packets. The experimental archive contains 300 information-ablation evaluations and 1,080 controlled scenario-condition runs generated from 60 base scenarios across six information-quality conditions and three replicates. These experiments provide internal computational verification and stress testing of the proposed framework; they are not field validation and should not be interpreted as observations from 1,080 independent highway projects. The repository includes the Python implementation, reproducible software-environment information, numerical verification scripts, experimental results, figures, metadata, licenses, checksums, and archived outputs from a frozen Grok 4.6 experiment comprising 60 ungrounded and 60 evidence-grounded completions. The deposited code loads and re-scores the archived LLM outputs; it does not regenerate hosted-model responses or query Grok. D2A-HWY-60 supports reproducible investigation of highway design-to-as-built correspondence, machine-readable construction acceptance rules, deterministic QA/QC decision support, PASS/NONCONFORMING/REVIEW REQUIRED decision logic, uncertainty-aware human review, evidence-grounded generative-AI reporting, and verified-as-built digital-twin information states. Numerical authority remains with the deterministic verification engine; the archived language-model experiment evaluates evidence-grounded reporting rather than autonomous construction acceptance. The benchmark uses research-constructed corridor models and synthetic as-built information. It does not contain live DOT project files, proprietary agency design files, field-survey measurements, or field-validation data. The primary S1 rule profile is based on publicly available UK Specification for Highway Works Series 700 provisions, while S2 provides specification-sensitivity testing using TxDOT Item 340 and documented research operationalisations. Gold labels are constructed using the S1 rule envelope; consequently, performance against those labels primarily evaluates implementation fidelity rather than independent field predictive validity. Researchers may reuse or extend the benchmark to evaluate alternative construction-tolerance profiles, geometric verification algorithms, information-quality and uncertainty policies, QA/QC decision architectures, or language models. The package is intended to support reproducibility and subsequent external validation using independently collected field data. Data license: CC BY 4.0.



