遇见数据集

K8s-Distill-Pilot: Corpus, Evaluation Outputs, LoRA Adapters, and Scripts for Context-Instrumental Data Distillation of Kubernetes Manifests

收藏
Zenodo2026-09-28 更新2026-10-01 收录
官方服务:

资源简介:

This record provides the corpus, fixed splits, teacher-generation batches with accepted and rejected outputs, student predictions, LoRA adapters, configurations, and pipeline scripts for the SpacSec 2026 paper "Context-Instrumental Data Distillation for Kubernetes Manifest Generation: Method and Experimental Evaluation." Training pairs were generated synthetically by DeepSeek-V4 Flash (API route deepseek-v4-flash) and retained only if they passed four validation levels: YAML syntax (L1), kubeconform --strict against Kubernetes 1.30.0 schemas (L2), two cross-resource checks (L3-lite), and critical Checkov policies (L4). The final deduplicated corpus contains 1,710 records; the reported split is train_1200 / validation_100 / test_200. Qwen2.5-Coder-1.5B-Instruct was fine-tuned with LoRA on CPU. The archive includes four evaluation runs on the fixed test_200 set (full-pass@1 of 164, 157, 182, and 183 out of 200), with raw and extracted predictions, per-record validation outcomes, and error audits. test_200 is an in-domain development benchmark, not an independent test on real-world manifests. The README documents the directory layout, record fields, validation levels, environment and validator versions, and reproduction steps. Base-model weights, the Kubernetes schema cache, and raw API response bodies are not included; request and response hashes are retained in every record. Licenses: data, evaluation outputs, and reports under CC BY 4.0; code under MIT; LoRA adapters under Apache 2.0, following the base model (see LICENSE-DATA.txt and LICENSE-CODE.txt).

提供机构:
Zenodo
创建时间:
2026-09-28
二维码
社区交流群
二维码
科研交流群
商业服务