Semantic Leakage from Compound Identifiers: Protocol, Data, and Code
收藏官方服务:
资源简介:
Pre-registered experiment testing whether the semantic content of compound identifiers built from meaningful words (e.g. {horse-farm-cow}) leaks into unrelated parts of a language model's output, and whether that leakage grows with exposure. Includes the frozen protocol, the run harness, verbatim raw call logs, derived scored tables, results, and the results memo. The upload is the complete archive tree at tag v1.0.0 of the GitHub repository (82 files; MANIFEST.sha256 covers the other 81 and is verified with "shasum -a 256 -c MANIFEST.sha256"). All content is synthetic model output; no human subjects. Data are licensed CC BY 4.0 and code MIT (LICENSE-data and LICENSE-code in the archive). Two dated deviations are recorded in amendments.md.
提供机构:
Zenodo创建时间:
2026-08-15



