Reproducibility Package for "A Governance Architecture for AI-Assisted Validation of Construction Cost Knowledge: Combining Domain Rules with LLM Semantic Analysis"
收藏资源简介:
This deposit contains the reproducibility materials for the manuscript "A Governance Architecture for AI-Assisted Validation of Construction Cost Knowledge: Combining Domain Rules with LLM Semantic Analysis." The work presents a deployed tender management system that validates Bill of Quantities (BOQ) pricing formulas at portfolio scale by combining trade-specific deterministic policy checks with large-language-model (LLM) semantic analysis under a four-layer hallucination-mitigation architecture (schema enforcement, evidence binding, deterministic override, read-only AI). Evaluation uses 17 construction tenders across two regions (Morocco, Ivory Coast), roughly 1.1 billion monetary units of portfolio value, and 100 expert-labelled formulas (inter-rater kappa = 0.96). Package contents (eight files): README.docx — Package overview, file listing, sanitization policy, and citation guidance. Reproducibility_Supplement.pdf — Full reproducibility documentation: dataset definitions, key counts, leakage prevention protocol, baseline model specifications, statistical analysis details, the complete verbatim LLM prompt template, and a worked input-output JSON trace for the REV-003 example. LLM_Prompt_Template_Full.docx — Complete verbatim system and user-message templates sent to the Azure OpenAI gpt-5.1-codex-mini deployment. Results.xlsx — Headline numbers from the manuscript in one workbook (Expert Metrics, Confusion Matrices, Cross-Region F1, Risk Distribution, Data Dictionary). Data/Expert_Validation_Sample.xlsx — The n=100 expert-labelled formula records. Data/Validation_Mode_Predictions.xlsx — Per-formula predictions for the four validation configurations (Hybrid, Rules-only, AI-only, AI-with-flags). Data/Baseline_Predictions.xlsx — Aggregated outputs for the six baseline models (LR-S2, XGB-S2, RF-S2, Isolation Forest, z-score, TF-IDF+LR), plus baseline ladder and fold composition. Data/Statistical_Tests_Frozen.xlsx — Frozen official statistical-test artefacts (McNemar, i.i.d. and TenderId-cluster bootstrap CIs, per-fold stability, inter-rater kappa). Sanitization: item descriptions, unit prices, and AI narrative text are blanked; FormulaId and TenderId are pseudonymized. Component percentages, expert labels, and risk classifications — the fields used by the analyses reported in the manuscript — are intact. All quantitative claims in the manuscript are computable from the released artefacts. Raw BOQ text and raw monetary amounts are not publicly available due to commercial confidentiality of active contractor pricing. Researchers requiring access to additional anonymized data may contact the corresponding author.



