遇见数据集

Self-Healing Cloud Infrastructure Systematic Review — Data Package

收藏
Zenodo2026-05-31 更新2026-05-26 收录
官方服务:

资源简介:

This deposit is the data layer for a PRISMA 2020 systematic review of machine learning approaches to self-healing cloud infrastructure. The corpus is 344 peer-reviewed and arXiv-indexed studies (2022–2026), identified from 1,537 unique post-deduplication records across seven academic databases (Web of Science was targeted in the protocol but inaccessible without institutional credentials). The deposit contains the canonical extraction dataset, the full screening trail, the methodology (locked rubrics, protocol, dated deviations log, PRISMA 2020 checklist), the synthesis tables and gap verdicts, the figures and their source code, the per-paper extraction Markdown files, and the reproducibility scripts. The accompanying journal manuscript is not part of this deposit; the publisher hosts the version of record. The canonical CSV (canonical_studies.csv) records per-study CASP item scores (Q1–Q6, 0–2 each), PROBAST domain ratings (Data Selection, Feature Engineering, Model Evaluation, Analysis Reporting; Low/High/Unclear with worst-domain rule for overall), inclusion rationale, primary phase classification (DETECT, DIAGNOSE, DECIDE, RECOVER, Enabling Tech, Cross-Phase Integration), ML technique with controlled-vocabulary and free-text detail, deployment context, evaluation methodology, dataset descriptor, and headline performance metrics. The dual-path sensitivity framework (full 344 corpus + 128-paper Low-ROB subset) can be reproduced by applying the named predicates in rq_to_answer_mapping.md to canonical_studies.csv. Two post-extraction audits are recorded in protocol_deviations.md: a duplicate-record audit collapsed 19 within-corpus duplicates missed at Phase 1 deduplication (audit_exclusion: dup_phase1), and a scope audit dropped 6 pre-2022 records retained in error after Refinement A (audit_exclusion: scope_refinement_a). The arithmetic 369 − 19 − 6 = 344 is enforced as an assertion in figures/prisma_flow.py and as a stale-value drift detector in audit/cross_doc_audit.py (61 claims, all reconcile to canonical_studies.csv). Methodology highlights:- Four-pass single-author screening (title, abstract, re-screen, full-text adjudication) with intra-rater reliability quantified via Cohen's κ and prevalence-adjusted bias-adjusted kappa (PABAK 0.905 three-category, exceeding the pre-specified 0.80 lock threshold).- CASP item-level scoring deposited per paper (casp_rubric.md).- PROBAST-adapted four-domain risk-of-bias scoring with worst-domain rule deposited per paper (probast_rubric.md).- Hybrid pipelines decomposed into nine architectural sub-classes (H1 LLM-orchestrated, H2 GNN+temporal, H3 Federated, H4 sequence+RL, H5 autoencoder+sequence, H6 metaheuristic+ML, H7 multi-modal fusion, H8 statistical+ML, H9 other multi-component) by a deterministic strict-priority classifier (hybrid_decomposition.md).- Manifest-driven cross-document numerical audit (audit/cross_doc_audit.py) plus stale-value drift detector against prior-version numbers. See README.md for a full file-by-file inventory.

提供机构:
Zenodo
创建时间:
2026-05-13
二维码
社区交流群
二维码
科研交流群
商业服务