遇见数据集

Does a Language Model Know When It's Done? — reproducibility package

收藏
Zenodo2026-07-13 更新2026-08-01 收录
官方服务:

资源简介:

# Does Termination Read From the Workspace? — Data & Code Release This is the reproducibility package for a J-space ablation experiment testing whether correct end-of-sequence (EOS) placement in an instruction-tuned language model depends on the "global workspace" representations identified by Anthropic's *Verbalizable Representations Form a Global Workspace in Language Models* (transformer-circuits.pub/2026/workspace/index.html). We ablate the Jacobian-lens-readable directions (condition B) against a norm-matched random-direction control (condition C) in `Qwen/Qwen2.5-1.5B-Instruct`, across a six-task battery designed to separate content difficulty from termination difficulty, and score stopping quality, generation fluency, and teacher-forced EOS calibration. Full method: `EXPERIMENT_PROTOCOL.md`. Full result: `RESULTS.md`. **Verdict:** both the k=32 main run and the k=8 sensitivity re-run land on the protocol's "ablation too aggressive" decision-table row — T1/T2 stopping is degraded in B vs C, but so is T6 fluency, so the stopping deficit cannot be formally dissociated from generic ablation damage at this model scale. Neither H1 nor H0 is established; see RESULTS.md §"Observations the verdict does not license as conclusions" for the (non-conclusive) trends and suggested follow-up at larger scale.

提供机构:
Zenodo
创建时间:
2026-07-10
二维码
社区交流群
二维码
科研交流群
商业服务