Data and scripts for: Same answers, different scores: an empirical study of evaluation choices in long-term memory benchmarks for LLM agents
收藏资源简介:
Raw data and scripts for the paper named in the title. The dataset holds saved runs of total-agent-memory 14.0.0 to 14.5.1 on LoCoMo, LongMemEval and MemoryAgentBench, cross-grading of TAM and Mem0 Platform answers by two judge configurations, per-question answers with judge verdicts, and the gate results of several retrieval and context changes. analysis.py recomputes every number in the paper from these files without API calls; its saved output is included. Third-party data: questions, gold answers and conversation excerpts from LoCoMo (CC BY-NC 4.0, non-commercial use only), LongMemEval (MIT) and MemoryAgentBench (MIT) appear in the answer files and stay under those original licenses; they are not covered by CC BY 4.0. README.md lists the files and the citations.



