Style Without Substance: Small-Model SFT Distillation of CTI-Style Prose and What It Reveals About Detector and Safety-Classifier Calibration (Paper 2 artefacts)
收藏资源简介:
This deposit accompanies the preprint "Cheap QLoRA distillation of attacker-persona prose from a security-tuned teacher: a substance-vs-form failure mode" (Paper 2). It contains the paper source, the eval harness, the six-axis attacker-persona rubric, per-record judgments under two judges (Claude Opus 4.7 and Claude Sonnet 4.6), an ATT&CK-validation output, a detector-baseline under three classifiers, and the 4,010-chunk public-CTI training corpus. Headline numbers reproducible from this deposit: Per-axis stylistic-fidelity means: 2.81/5 under Opus 4.7, 2.38/5 under Sonnet 4.6, per-record Pearson r = 0.806 (N=74). ATT&CK technique-to-ID confabulation: 35/36 = 97.2% on name-paired citations, validated against the MITRE Enterprise matrix. Detector TPR collapse on the ttp_summary transform: 20% under Opus and Gemma 4 E4B, 40% under qwen2.5:7b, vs 60-100% on the other two transforms. The trained QLoRA adapter is NOT included in this deposit. The substance-vs-form gap means the adapter is a competent attacker-prose generator with low marginal scientific value beyond what the eval JSONLs already demonstrate; we judge the dual-use cost-benefit unfavourable for open release. The training recipe in scripts/ plus the public-CTI corpus in training_corpus/ are sufficient to reproduce the adapter on a single 16 GB GPU in ~5.5 hours. Reuse: CC-BY-4.0 over the entire deposit. See README.md inside the zip for the bundle layout, file-by-file reproduction recipe, and a diagram of how the harness fits together. Cite the deposit DOI together with the arxiv ID of the preprint when reusing any artefact. Companion deposits: Paper 1 (LLM-targeted recon honeypot measurement) and the safety-classifier note (Opus-specific format-driven refusal) each get their own Zenodo DOIs; cross-links added on publication.



