Run artifacts for "Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings" (arXiv:2607.21962)
收藏资源简介:
Complete judged-verdict and run artifacts for arXiv:2607.21962. Every headline figure in the paper recomputes from these files by arithmetic alone — no provider calls and no cost. Verified before deposit: extracting this archive and applying the deduplication recipe in runs/README.md reproduces the published cross-family confusion table exactly — 8,985 answers as 6,827 / 1,029 / 54 / 1,075, agreement 87.9%, κ = 0.600. Read runs/README.md first. fullhistory was re-judged twice and the archive retains both passes; they disagree on 131 of 1,797 answers (7.3%). crossjudge_full/ is canonical. The duplicate is kept rather than cleaned up because it is an accidental test–retest measurement of the configured judge service, and deduplicating the release would destroy it. Not included: the LMSYS corpus (gated and non-redistributable under its licence) and the LongMemEval extraction cache (derivative of benchmark conversations, treated the same way). Neither backs any headline claim in the paper. Contents: s0/ s1/ s2/ per-seed primary-judge verdicts; baselines/ fullhistory and token-matched window arms; tenure/ longitudinal checkpoints; crossjudge_full/ canonical cross-family re-judge; crossjudge_vector/ the second pass; security/ injection-probe results; emb_cache.jsonl hash-keyed embedding vectors (no text); llm_calls.jsonl call metrics only (no prompts). 455 files, 54 MB uncompressed. Per-file SHA-256 checksums are in the accompanying MANIFEST.sha256.



