RULER Synthetic-Task Model Outputs and Score Archives for "Do Short-Context Mechanistic Probes Predict Long-Context Capability?"
收藏资源简介:
Synthetic-task raw prediction archives and full-task score summaries for the paper "Do Short-Context Mechanistic Probes Predict Long-Context Capability?" — 13 RULER measurement runs (Gemma-3 270m/1b/4b/12b/27b at up to 128K context, Qwen3 0.6b/1.7b; 11 of the 13 runs use the patched `answer_prefix` protocol, and the other two — a gemma-3-270m and a qwen3-0.6b run — repeat those two models under the pre-fix condition), 500 examples per task, all measured context lengths. Harness: NVIDIA/RULER at commit `38da79d79519ef87aa46ae804f838e1eab7f86d7`, with the one-line `answer_prefix` re-attachment fix (upstream issue #107 / PR #108). This record releases, per archive, the six RULER tasks whose prompts are fully synthetic (`niah_single_1`, `niah_multikey_2`, `niah_multikey_3`, `vt`, `cwe`, `fwe` — these tasks' prompts are generated from word lists, numbers, and noise sentences, not from any third-party corpus) as byte-for-byte JSONL rows, plus the official evaluator's regenerated 6-task `summary.csv`/`submission.csv` for each of the 40 measured cells (model x context length). Outputs of the seven tasks whose prompts embed third-party source text (five needle-in-a-haystack tasks over Paul Graham essays; SQuAD and HotpotQA passages, each CC BY-SA 4.0) are withheld in their entirety from this record — no `input`, `pred`, or `outputs` field of those tasks appears here. Aggregate echo-rate statistics for the withheld tasks are published in `REDACTION_NOTE.md`, together with the reasoning for withholding rather than releasing a masked version (a masking rule was evaluated separately and found score-preserving, but a more conservative release scope was chosen); per-record SHA-256 commitments for the withheld rows are preserved in a private audit-grade layer whose archive-level checksums are published in `SHA256SUMS_AUDIT.txt`. `score_summaries.tar.gz` additionally provides the full 13-task `summary.csv` for every one of the 40 measured cells (scores and null counts only, no prompt or prediction text) — every RULER score, null count, and effective-context-length input that the paper reports for these 13 runs derives from this file, and it is byte-identical to the measurement-time summaries (verified; per-cell gate results in `synthetic_release_report.json`). Effective lengths for models outside these runs are recomputed by the paper from publicly reported score tables, and the paper's other quantities (capacity-lens statistics, probe measurements, and the regressions computed from them) come from separate measurement artifacts not part of this record. Reconstruction of the withheld tasks' prompts is deterministic from the pinned harness commit, fixed seeds, the model tokenizer, and the source-material checksums in `MATERIALS_SHA256.txt`, and was verified end-to-end on 2026-07-21 (18,000/18,000 SHA-256 matches on live re-download and re-generation; recipe and entry-point scripts in `REDACTION_NOTE.md`, `reconstruct_inputs.sh`, `verify_reconstruction.py`). The audit-grade intermediate layer and the unredacted originals this record derives from are preserved by the author; their SHA-256 checksums are published in `SHA256SUMS_AUDIT.txt` and `SHA256SUMS_FULL_PRIVATE.txt`. License scope: the CC BY 4.0 license of this record covers the author-created content released here — the six synthetic tasks' prediction records, the full-summary scores, audit metadata, tooling, and documentation. No released rows originate from tasks embedding the identified third-party corpora (the six released tasks' prompts are generated from word lists, numbers, and noise sentences). No exact matches of ≥60 Unicode code points were detected against any of the three corpora in any released field — this is gate G3 of this record, whose result (360 files scanned, zero hits) is recorded in `synthetic_release_report.json`; the scan was run twice against the same released bytes by two independently written implementations. A separate threshold-sensitivity run, preserved by the author (its result table is not part of this record; available on request), additionally found no match of ≥40 code points against the Paul Graham corpus in any of the 120,000 released predictions. As with any raw LLM output, the absence of memorized text from sources outside the three scanned corpora cannot be exhaustively proven; the released prompts embed no third-party corpus by construction. RULER harness (Apache-2.0, https://github.com/NVIDIA/RULER): task templates and `answer_prefix` strings quoted under its license.



