Edge LLM and agentic AI in BACS: scoping review corpus and pipeline
收藏资源简介:
Data and code deposit accompanying the scoping review “Edge language models and agentic AI in building automation: a scoping review of evidence maturity and authority tiers” (Simson, Kiil, Võsa & Kurnitski, 2026). Contents: Full deduplicated corpus with screening decisions (S1, 1,281 records). In-window themed corpus (S2, 272 rows) with A1–A10 theme tags and rule-based E1–E4 / T1–T4 heuristic columns. S2 retains the full audit trail: 266 final_status = INCLUDE, 2 BACKGROUND, and 4 document-type EXCLUDE rows; after post-coding scope review these correspond to 257 INCLUDE, 5 INCLUDE_NON_LM contextual comparators, 4 OUT_OF_SCOPE_RECLASSIFIED, 2 BACKGROUND, and 4 EXCLUDE rows. Query log: 27 Boolean queries, final search 30 May 2026, plus 12 pre-2022 backward-citation records. Per-record manual E-tier and T-tier coding sheet for all 266 records that passed the screening pipeline before post-coding scope review (S3). Filtering S3 by scope_status = INCLUDE gives the 257 LM/agentic-AI primary records used for Table 2 of the main text; filtering by scope_status in {INCLUDE, INCLUDE_NON_LM} gives the 262 in-scope primary records. Cross-study quantitative extraction (S4, 23 named studies). Dual-screener audit on a 60-record stratified sample (Cohen’s κ = 0.733–0.895 across screener/pipeline comparisons), a 50-record AUTO_EXCLUDE recall check with 0 unambiguous false negatives, and a 58-record second-coder E×T reliability audit (κE = 0.786, κT = 0.715). Targeted preprint/conference scan covering ACM BuildSys, ACM e-Energy, arXiv, IEEE smart-building / smart-grid proceedings, and 20 LLM-BACS candidates. Consolidated Python screening pipeline and seven-stage abstract-retrieval scripts: Crossref → OpenAlex → Semantic Scholar → Springer Nature Meta → Elsevier ScienceDirect → Elsevier Scopus → Google Scholar via scholarly. E1–E4 / T1–T4 coding rubric with 12 borderline-case rules. PRISMA-ScR checklist and extended narrative supplements, including the cybersecurity threat model, ten-gap research roadmap, limitations section, and per-application extended synthesis. Supplementary/source figures S1 and S2: Sankey theme × year diagram and journal × year heatmap, provided as PNG and SVG. Licence: CC-BY-4.0 for data and documents; MIT for code. See LICENSE-CC-BY-4.0.txt and LICENSE-MIT.txt in the deposit. Publisher API access details are provided only as credential placeholders and configuration notes. Use of the Springer Nature and Elsevier APIs remains subject to the respective publisher API terms of service and is not covered by the MIT licence.
本数据集与代码附件配套于范围综述(scoping review)“边缘语言模型与AI智能体(AI Agent)在建筑自动化中的应用:证据成熟度与权威层级的范围综述”(Simson、Kiil、Võsa与Kurnitski,2026)。 内容: 完整去重语料库(deduplicated corpus)与筛选决策数据集(S1,含1281条记录)。 窗口内主题语料库(S2,272条数据行),搭载A1–A10主题标签(theme tags)与基于规则的E1–E4、T1–T4启发式列(heuristic columns)。S2保留完整审计轨迹(audit trail):266条final_status字段值为INCLUDE(纳入)的记录、2条为BACKGROUND(背景)的记录与4条document-type为EXCLUDE(排除)的记录;经后编码范围审查后,这些记录被调整为257条INCLUDE(纳入)、5条INCLUDE_NON_LM(非大语言模型纳入)上下文对比记录、4条OUT_OF_SCOPE_RECLASSIFIED(范围外重分类)记录、2条BACKGROUND(背景)记录与4条EXCLUDE(排除)记录。 查询日志:共27个布尔查询(Boolean queries),最终检索时间为2026年5月30日,另包含12条2022年前的回溯引用记录(backward-citation records)。 针对后编码范围审查前通过筛选流程(screening pipeline)的全部266条记录,提供逐记录手动E层级(E-tier)与T层级(T-tier)编码表(coding sheet)(S3)。通过以scope_status = INCLUDE(纳入)对S3进行过滤,可得到主文表2所用的257条大语言模型(LLM)/AI智能体核心记录;通过以scope_status ∈ {INCLUDE, INCLUDE_NON_LM}进行过滤,可得到262条纳入范围的核心记录。 跨研究定量提取数据集(S4,含23项已标注研究)。 针对60条分层抽样样本(stratified sample)开展双筛选员审计,其筛选员/流程对比的科恩κ系数(Cohen’s κ)为0.733–0.895;另包含50条AUTO_EXCLUDE(自动排除)召回检查样本(无明确假阴性(false negatives)结果),以及58条二次编码员E×T信度审计(reliability audit)样本,其中κE=0.786,κT=0.715。 针对性预印本(preprint)/会议检索覆盖ACM BuildSys、ACM e-Energy、arXiv、IEEE智能建筑/智能电网会议论文集(proceedings),以及20个大语言模型(LLM)-BACS候选对象。 整合后的Python筛选流程与七阶段摘要检索脚本:Crossref → OpenAlex → Semantic Scholar → Springer Nature Meta → Elsevier ScienceDirect → Elsevier Scopus → 基于scholarly库的Google Scholar检索。 包含12条边界案例规则(borderline-case rules)的E1–E4、T1–T4编码准则(coding rubric)。 PRISMA-ScR检查表与扩展叙事补充材料,涵盖网络安全威胁模型、十差距研究路线图、局限性章节,以及逐应用场景的扩展综合分析。 补充/源图表S1与S2:分别为桑基主题×年份图与期刊×年份热图,以PNG和SVG格式提供。 许可协议:本数据集与文档采用CC-BY-4.0许可协议,代码采用MIT许可协议。详见附件中的LICENSE-CC-BY-4.0.txt与LICENSE-MIT.txt。出版商API访问详情仅以凭证占位符与配置说明形式提供。使用Springer Nature与Elsevier API仍需遵守对应出版商的API服务条款,不受MIT许可协议覆盖。



