WizeMe Memory Benchmark Receipts: LoCoMo QA and Retrieval, LongMemEval Retrieval, June 2026
收藏资源简介:
Public benchmark receipts for WizeMe memory work, including LoCoMo and LongMemEval retrieval metrics plus a scoped 300-question LoCoMo end-to-end QA result generated from the current code path. Retrieval and end-to-end QA are reported as separate metric families. The QA result is not a full-dataset or same-mode provider superiority claim. The package is intentionally limited to public receipts and excludes private source code, user data, training data, secrets, and other non-public materials. LoCoMo and LongMemEval use different evaluation protocols. LoCoMo Any@3 measures exact-turn retrieval across tightly clustered sessions. LongMemEval Any@3 measures answer-cluster retrieval across a larger haystack. Both are reported raw without cross-benchmark normalization.



