遇见数据集

BACS LLM / agentic-AI scoping review (2022–2026): corpus, screening pipeline, and dual-screener audit

收藏
Zenodo2026-06-01 更新2026-06-05 收录
官方服务:

资源简介:

Version 1.1.0 (2026-05-31). See CHANGELOG.md in the deposit for the full v1.0 → v1.1 difference list. Companion data and code deposit for the scoping review “Edge language models and agentic AI in building automation: a scoping review”. The deposit is journal-agnostic: it is the canonical citation for the underlying corpus, screening decisions, and audit artefacts irrespective of where the article is eventually published. This deposit contains the machine-readable corpus, the consolidated nine-step Python screening/classification pipeline, the dual-screener audit artefacts (60-record stratified sample, two independent screener decision files, and the Cohen’s κ computation script), the supplementary narrative documents, the PRISMA-ScR checklist, and the supplementary figures. The deposit supports downstream verification of every quantitative claim in the companion manuscript and partial reproduction of the corpus from the deposited CSV outputs. The article DOI will be added as an “is supplement to” related identifier on this Zenodo record once the companion article is accepted (metadata-only update; no version bump). Headline figures reproducible from this deposit: 9,834 cumulative search returns → 1,281 deduplicated records → 262 primary INCLUDE + 2 BACKGROUND methodology reviews = 264 in-window journal records (January 2022 – May 2026). 27 Boolean queries across six sources: Scopus, ScienceDirect, Semantic Scholar via Wiley discovery wrapper, Consensus, PubMed, and Tavily. Dual-screener Cohen’s κ: between-screener κ = 0.895; screener-vs-pipeline κ = 0.733–0.833, indicating substantial agreement. Reproducible by running dual_screening/compute_kappa.py. 78-paper top-cited extraction with the manual coding underlying Table 2.1 of the manuscript. Heuristic columns are deposited; the per-record manual coding sheet is currently a residual reproducibility gap noted in §7.1, limitation 6 of the manuscript. 23 named studies with manually coded evidence-tier / authority-tier classification in S4_Cross_study_quantitative.csv. The deposit does not include the manuscript itself (copyright held by the journal under standard publishing agreements) or the main-text figures embedded in the article. See README.md inside the bundle for the full data dictionary and the legacy T1–T10 ↔ A1–A10 theme-prefix note. Licences: CC-BY-4.0 for data and documents; MIT for code. See LICENSE inside the bundle for details.

版本1.1.0(2026年5月31日)。如需了解v1.0至v1.1的完整变更详情,请查阅存档包内的CHANGELOG.md文件。 本存档为范围综述《建筑自动化中的边缘语言模型与AI智能体(AI Agent):一项范围综述》的配套数据与代码存档。本存档不受期刊限制:无论最终文章发表于何处,本存档均为底层语料库、筛选决策与审核工件的标准引用源。 本存档包含可机读语料库、整合式九步Python筛选/分类流水线、双审稿人审核工件(60条记录的分层样本、两份独立审稿人决策文件,以及科恩κ系数(Cohen’s κ)计算脚本)、补充叙述性文档、PRISMA-ScR清单,以及补充图表。本存档可支持对配套稿件中所有定量结论的下游验证,以及通过存档的CSV输出结果对语料库进行部分复现。待配套稿件被录用后,本Zenodo记录将添加标注为“作为补充”的相关标识符以关联文章DOI(仅更新元数据,不触发版本号升级)。 可通过本存档复现的核心图表如下: 累计检索结果9834条 → 去重后得到1281条记录 → 262条优先纳入的主文献 + 2篇背景方法学综述 = 264条符合时间窗口(2022年1月—2026年5月)的期刊文献记录。 共27个布尔查询式,覆盖6个数据源:Scopus、ScienceDirect、通过Wiley发现接口调用的Semantic Scholar、Consensus、PubMed以及Tavily。 双审稿人科恩κ系数:审稿人间κ值为0.895;审稿人与流水线间κ值为0.733~0.833,表明实质性一致。可通过运行dual_screening/compute_kappa.py复现该结果。 本存档包含78篇高被引论文的提取数据,其手动编码结果对应稿件中表2.1的内容。启发式列已存入存档;每条记录的手动编码表目前存在一处复现性缺口,详见稿件第7.1节局限性6。 S4_Cross_study_quantitative.csv中包含23项已命名研究,其证据层级/权威层级分类均为手动编码所得。 本存档不包含稿件本身(依据标准出版协议,稿件版权归期刊所有)以及文章内嵌的正文中的图表。如需完整数据字典以及旧版T1–T10 ↔ A1–A10主题前缀说明,请查阅存档包内的README.md文件。 许可协议:数据与文档采用CC-BY-4.0协议;代码采用MIT协议。详情请查阅存档包内的LICENSE文件。

提供机构:
Zenodo
创建时间:
2026-06-01
二维码
社区交流群
二维码
科研交流群
商业服务