Measurement artifacts for "The Checkpoint Grid and the Reference Point: Measuring Capability Retention Along Low-Data QLoRA Trajectories" (ICITEE 2026)
收藏资源简介:
Evaluation records, frozen analysis contract, and analysis code for the ICITEE 2026 paper "The Checkpoint Grid and the Reference Point: Measuring Capability Retention Along Low-Data QLoRA Trajectories". Typhoon-1.5-8B-Instruct was adapted with QLoRA on 80 Thai or Vietnamese SeaBench examples over a 500-step horizon and three training seeds. The primary training trajectories are held constant while their measurement varies — the checkpoint grid, the in-context exemplar draw, the aggregate, the label-scoring rule, and the reference point against which deterioration is declared. The deposit contains the machine-readable analysis contract frozen before any reported run was trained; per-item XNLI predictions and held-out likelihoods for the paired analyses, plus aggregate records for the remaining evaluated checkpoints; training configurations and prompt specifications; and the five reported robustness and sensitivity analyses — a 16-draw exemplar campaign with its draw manifest, a warm-up-schedule control over {10, 50, 100} steps at three seeds (72 evaluation cells), a checkpoint-selection resampling analysis, a FLORES-200 Thai language-modelling measure, and a six-permutation counterbalanced-label probe — together with the CPU-only analysis code that reproduces every reported number without a GPU. Only the warm-up control changes the training schedule. Neither XNLI nor SeaBench source text is redistributed: exemplar and item identities appear as indices and SHA-256 hashes. The manuscript itself is not deposited. Cluster filesystem paths have been replaced with neutral placeholders, tabulated in SCRUB_MAP.md; no score, prediction, interval, or identity was altered. Every extracted payload file except MANIFEST.json itself carries a SHA-256 in that manifest; the seven upload archives were also checked before upload.



