Reproducibility Package for "A Governance Architecture for AI-Assisted Validation of Construction Cost Knowledge: Combining Domain Rules with LLM Semantic Analysis"
收藏资源简介:
This deposit contains the reproducibility materials for the manuscript "A Governance Architecture for AI-Assisted Validation of Construction Cost Knowledge: Combining Domain Rules with LLM Semantic Analysis." The work presents a deployed tender management system that validates Bill of Quantities (BOQ) pricing formulas at portfolio scale by combining trade-specific deterministic policy checks with large-language-model (LLM) semantic analysis under a four-layer hallucination-mitigation architecture (schema enforcement, evidence binding, deterministic override, read-only AI). Evaluation uses 17 construction tenders across two regions (Morocco, Ivory Coast), roughly 1.1 billion monetary units of portfolio value, and 100 expert-labelled formulas (inter-rater kappa = 0.96). Package contents (eight files): README.docx — Package overview, file listing, sanitization policy, and citation guidance. Reproducibility_Supplement.pdf — Full reproducibility documentation: dataset definitions, key counts, leakage prevention protocol, baseline model specifications, statistical analysis details, the complete verbatim LLM prompt template, and a worked input-output JSON trace for the REV-003 example. LLM_Prompt_Template_Full.docx — Complete verbatim system and user-message templates sent to the Azure OpenAI gpt-5.1-codex-mini deployment. Results.xlsx — Headline numbers from the manuscript in one workbook (Expert Metrics, Confusion Matrices, Cross-Region F1, Risk Distribution, Data Dictionary). Data/Expert_Validation_Sample.xlsx — The n=100 expert-labelled formula records. Data/Validation_Mode_Predictions.xlsx — Per-formula predictions for the four validation configurations (Hybrid, Rules-only, AI-only, AI-with-flags). Data/Baseline_Predictions.xlsx — Aggregated outputs for the six baseline models (LR-S2, XGB-S2, RF-S2, Isolation Forest, z-score, TF-IDF+LR), plus baseline ladder and fold composition. Data/Statistical_Tests_Frozen.xlsx — Frozen official statistical-test artefacts (McNemar, i.i.d. and TenderId-cluster bootstrap CIs, per-fold stability, inter-rater kappa). Sanitization: item descriptions, unit prices, and AI narrative text are blanked; FormulaId and TenderId are pseudonymized. Component percentages, expert labels, and risk classifications — the fields used by the analyses reported in the manuscript — are intact. All quantitative claims in the manuscript are computable from the released artefacts. Raw BOQ text and raw monetary amounts are not publicly available due to commercial confidentiality of active contractor pricing. Researchers requiring access to additional anonymized data may contact the corresponding author.
本数据集包含论文《面向建筑造价知识AI辅助校验的治理架构:结合领域规则与大语言模型语义分析》的可复现研究材料。 本研究构建了一套已部署的招标管理系统,可在项目组合规模下校验工程量清单(Bill of Quantities, BOQ)的计价公式,核心逻辑为将针对细分行业的确定性策略校验,与四层幻觉缓解架构(模式强制、证据绑定、确定性覆盖、只读AI)下的大语言模型(Large Language Model, LLM)语义分析相结合。 本次评估共使用来自两个地区(摩洛哥、科特迪瓦)的17份建筑招标项目,覆盖约11亿货币单位的项目组合价值,以及100份经专家标注的计价公式(评分者间κ系数=0.96)。 本数据包包含8个文件: 1. README.docx — 数据包概述、文件清单、脱敏规则与引用指南 2. Reproducibility_Supplement.pdf — 完整可复现性文档:涵盖数据集定义、核心统计量、防泄露协议、基线模型规格、统计分析细节、大语言模型完整逐字提示模板,以及REV-003示例的完整输入输出JSON追踪记录 3. LLM_Prompt_Template_Full.docx — 发送至Azure OpenAI gpt-5.1-codex-mini部署环境的完整逐字系统提示与用户消息模板 4. Results.xlsx — 承载论文核心数据的工作簿,包含专家指标、混淆矩阵、跨区域F1分数、风险分布、数据字典 5. Data/Expert_Validation_Sample.xlsx — 包含100份经专家标注的计价公式记录的数据集(n=100) 6. Data/Validation_Mode_Predictions.xlsx — 四种校验配置(混合模式、仅规则模式、仅AI模式、带标记AI模式)的单条公式预测结果 7. Data/Baseline_Predictions.xlsx — 六种基线模型的聚合输出结果(LR-S2、XGB-S2、RF-S2、孤立森林、z分数、词频-逆文档频率+LR(TF-IDF+LR)),以及基线模型层级与折次构成 8. Data/Statistical_Tests_Frozen.xlsx — 固化的正式统计检验产物(McNemar检验、独立同分布与TenderId聚类自助法置信区间、单折稳定性、评分者间κ系数) 脱敏说明:项目描述、单价与AI叙事文本均已做空白处理;FormulaId与TenderId已完成假名化。而本研究分析所用的字段(成分占比、专家标注结果与风险分类)均保留完整。论文中所有量化结论均可通过本次发布的研究产物复现计算。 由于活跃承包商的计价数据涉及商业机密,原始工程量清单文本与原始货币金额未对外公开。如需获取额外匿名化数据的研究人员,请联系通讯作者。



