遇见数据集

Style Without Substance: Small-Model SFT Distillation of CTI-Style Prose and What It Reveals About Detector and Safety-Classifier Calibration (Paper 2 artefacts)

收藏
Zenodo2026-05-25 更新2026-05-26 收录
官方服务:

资源简介:

This deposit accompanies the preprint "Style Without Substance: Small-Model SFT Distillation of CTI-Style Prose and What It Reveals About Detector and Safety-Classifier Calibration" (Paper 2). It contains the paper source, the eval harness, the six-axis attacker-persona rubric, per-record judgments under three Opus passes (original N=90, re-judged 15 lost records, and a fresh N=194 wider eval), a Sonnet 4.6 inter-rater pass on N=90, a Gemma 4 E4B inter-rater pass on N=15, an ATT&CK-validation output, a detector-baseline under three classifiers, and the 4,010-chunk public-CTI training corpus. v2 corrects a v1 methodological error. The v1 deposit (10.5281/zenodo.20365823) reported a 53% Opus-refusal rate on the ttp_summary transform class. Re-investigation on 2026-05-24 established that 15 of the 16 originally-reported refusals were upstream BadRequestError: credit balance too low events that returned before the request reached Opus; only 1/30 was a genuine refusal (chunk_idx=2 ttp_summary; reproduces at 1/67 on the new N=194 cut, same record both times). v2 includes the corrected eval data, a new N=194 wider eval that tightens per-axis CIs from ±0.13-0.29 to ±0.086-0.170, and the v3-vs-v4 audit trail for the companion safety-note. Headline numbers reproducible from this v2 deposit: Per-axis Opus 4.7 stylistic-fidelity means on N=194 wider eval: overall 2.76/5, style_consistency 3.46 (only axis above acceptable) down to tradecraft_vocabulary 2.33. CI half-widths ±0.086 to ±0.170. Two-judge inter-rater on N=89 corrected cut: Opus 2.72, Sonnet 2.38, per-record Pearson r = 0.814, mean shift -0.39. Bidirectional calibration bracket with Gemma +0.42. ATT&CK technique-to-ID confabulation: 35/36 = 97.2% on name-paired citations, validated against the MITRE Enterprise matrix. Detector TPR collapse on the ttp_summary transform: 20% under Opus and Gemma 4 E4B, 40% under qwen2.5:7b, vs 60-100% on the other two transforms. Opus 4.7 single-record content-specific refusal: 1/30 on N=90, 1/67 on N=194, same chunk_idx=2 ttp_summary record both times. The record falsely attributes attacker-controlled distribution channels to real US-CERT / CISA infrastructure; seven other judges (Sonnet 4.6, Haiku 4.5, Gemma 4 E4B, foundation-sec, qwen2.5:7b, gpt-oss-120b, Llama-4-Scout) all engage. Full analysis in the companion safety-note v4. The trained QLoRA adapter is NOT included in this deposit. The substance-vs-form gap means the adapter is a competent attacker-prose generator with low marginal scientific value beyond what the eval JSONLs already demonstrate; we judge the dual-use cost-benefit unfavourable for open release. The training recipe in scripts/ plus the public-CTI corpus in training_corpus/ are sufficient to reproduce the adapter on a single 16 GB GPU in ~5.5 hours. Reuse: CC-BY-4.0 over the entire deposit. See README.md inside the zip for the bundle layout, file-by-file reproduction recipe, and a diagram of how the harness fits together. Cite the deposit DOI together with the arxiv ID of the preprint when reusing any artefact. Companion deposits: Paper 1 (LLM-targeted recon honeypot measurement) and the safety-classifier note (Opus-specific single-record content guardrail) each get their own Zenodo DOIs; cross-links added on publication.

提供机构:
Zenodo
创建时间:
2026-05-24
二维码
社区交流群
二维码
科研交流群
商业服务