K8s-Distill-Pilot: Corpus, Evaluation Outputs, LoRA Adapters, and Scripts for Context-Instrumental Data Distillation of Kubernetes Manifests
收藏资源简介:
This record provides the corpus, fixed splits, teacher-generation batches with accepted and rejected outputs, student predictions, LoRA adapters, configurations, and pipeline scripts for the SpacSec 2026 paper "Context-Instrumental Data Distillation for Kubernetes Manifest Generation: Method and Experimental Evaluation." Training pairs were generated synthetically by DeepSeek-V4 Flash (API route deepseek-v4-flash) and retained only if they passed four validation levels: YAML syntax (L1), kubeconform --strict against Kubernetes 1.30.0 schemas (L2), two cross-resource checks (L3-lite), and critical Checkov policies (L4). The final deduplicated corpus contains 1,710 records; the reported split is train_1200 / validation_100 / test_200. Qwen2.5-Coder-1.5B-Instruct was fine-tuned with LoRA on CPU. The archive includes four evaluation runs on the fixed test_200 set (full-pass@1 of 164, 157, 182, and 183 out of 200), with raw and extracted predictions, per-record validation outcomes, and error audits. test_200 is an in-domain development benchmark, not an independent test on real-world manifests. The README documents the directory layout, record fields, validation levels, environment and validator versions, and reproduction steps. Base-model weights, the Kubernetes schema cache, and raw API response bodies are not included; request and response hashes are retained in every record. Licenses: data, evaluation outputs, and reports under CC BY 4.0; code under MIT; LoRA adapters under Apache 2.0, following the base model (see LICENSE-DATA.txt and LICENSE-CODE.txt).



