From Active Learning to Inference-Time Governance for LLM-Assisted Code Generation: A Systematic Literature Review with a Review-Derived PLC Governance Framework
收藏资源简介:
This document serves as the companion reproducibility package for themanuscript\emph{From Active Learning to Inference-Time Governance for LLM-AssistedCode Generation: A Systematic Literature Review with a Review-DerivedPLC Governance Framework}.Its purpose is to provide a transparent and auditable record of thesearch, screening, eligibility assessment, quality appraisal,record-level reconciliation, and repeatability analyses used to constructthe formal evidence base of the systematic literature review. The review is organized across three linked evidence layers:\textbf{(i)} Active Learning foundations and AL-relevant general machinelearning;\textbf{(ii)} inference-time mechanisms for LLM-assisted executable-codegeneration, completion, repair, testing, verification, and validation; and\textbf{(iii)} AL-relevant evidence for PLC-oriented code generation,synthesis, verification, and safety-critical control.The systematic review constitutes the primary scientific contribution.The resulting PLC governance framework is a downstream, review-deriveddesign roadmap for future implementation and empirical evaluation and isnot treated as an independently validated PLC toolchain. Two complementary formal retrieval routes were used.The principal broad-index searches of\href{https://www.scopus.com/search/form.uri?display=advanced}{Scopus}and\href{https://ieeexplore.ieee.org/search/advanced/command}{IEEE Xplore}produced \textbf{711 pre-deduplication records}, comprising\textbf{386 Scopus records} and \textbf{325 IEEE Xplore records}.Complementary venue-targeted searches of the\textbf{ACM Digital Library}, \textbf{ACL Anthology},\textbf{official ICLR proceedings}, and\textbf{official NeurIPS proceedings} were used to identifyuncertainty-aware LLM code-generation studies that may not use explicitActive-Learning terminology. The two retrieval routes were maintained separately for sourceprovenance. Venue-targeted records were assessed using the samepublication-status, eligibility, primary-layer-assignment, andquality-appraisal criteria and were reconciled against the broad-indexevidence base before construction of the final formal corpus.The primary Scopus/IEEE Xplore searches and broad-index LLM--PLC updatewere completed on \textbf{15 July 2026}, while the venue-targeted searcheswere completed on \textbf{2 August 2026}, which defines the upper temporalboundary of the review. Within the principal broad-index screening stream,\textbf{76 duplicate records} were removed, leaving\textbf{635 records} for title--abstract screening.Of these, \textbf{280 records} were excluded and\textbf{355 reports} were sought for full-text retrieval.After \textbf{9 reports} could not be retrieved,\textbf{346 full-text reports} were assessed. Following eligibility assessment, quality appraisal, and record-levelreconciliation, \textbf{223 broad-index reports were not retained},leaving \textbf{123 unique broad-index studies} in the formal evidencebase. The complementary venue-targeted route contributed\textbf{5 additional unique formal studies} after reconciliation againstthe broad-index evidence base: \textbf{2 ACL Anthology studies} toLayer~(ii), and \textbf{2 ACM Digital Library studies} plus\textbf{1 ACL Anthology study} to Layer~(iii). The final reconciled formal corpus therefore comprises\textbf{128 unique studies}:\textbf{60 studies in Layer~(i)},\textbf{25 studies in Layer~(ii)}, and\textbf{43 studies in Layer~(iii)}.Each canonical study is assigned to one primary evidence layer and countedonce in the formal corpus. Formal evidence is restricted to verified peer-reviewed journal articlesand accepted peer-reviewed conference papers processed through thedocumented review workflow. Preprints, tools, unpublished materials, andother sources outside this workflow are maintained separately assupplementary contextual evidence and do not contribute to the formalstudy-selection counts, formal evidence maps, research-questionconclusions, cross-layer findings, or review-derived frameworkrequirements. In addition to the search and selection audit trail, this companiondocuments the reproducibility of the single-reviewer assessment procedure.\textbf{Section~12} provides the record-level derivation of theintra-rater screening-repeatability results reported in manuscriptTable~A.12, including title--abstract screening, full-text eligibility,primary-layer assignment, Cohen's \(\kappa\), and bootstrap confidenceintervals.\textbf{Section~13} provides the corresponding derivation of thequality-appraisal repeatability results reported in manuscriptTable~A.13, including the five ordinal appraisal domains, linearlyweighted Cohen's \(\kappa\), the total-score \(\mathrm{ICC(A,1)}\),bootstrap confidence intervals, data-integrity checks, and thereproduction workflow. Overall, this document is intended to function as a standalonereproducibility and audit record for the systematic-review component ofthe study. It enables an independent researcher to reconstruct the searchlogic, verify the PRISMA and corpus counts, trace record-level provenanceand appraisal decisions, and reproduce the screening and quality-appraisalrepeatability analyses. These materials support reproducibility of thesystematic review and its subsequent evidence-to-design derivation; theydo not constitute implementation-level evidence or empirical validationof the downstream PLC governance framework.



