WizeMe Memory Benchmark Receipts: LoCoMo QA and Retrieval, LongMemEval Retrieval, June 2026
收藏资源简介:
Public benchmark receipts for WizeMe memory work, including LoCoMo and LongMemEval retrieval metrics plus scoped 300-question LoCoMo end-to-end QA policy and stability receipts. The lead LoCoMo QA number is the current 76.06% policy score, and that policy score is official-style 1:1. Legacy internal pass-rate diagnostics are excluded from public/provider-comparable scoring. The older 83.67% three-run all-category result is retained as a historical stability control, not the current headline. Conservative quality-lane answer p95 is 14.630 seconds, while the faster lane records 7.047 seconds. These seconds are model-generation and judging latency, not retrieval latency. Retrieval and end-to-end QA are reported as separate metric families. The QA result is not a full-dataset or same-mode provider superiority claim. The package is intentionally limited to public receipts and excludes private source code, user data, training data, secrets, and other non-public materials. LoCoMo and LongMemEval use different evaluation protocols. LoCoMo Any@3 measures exact-turn retrieval across tightly clustered sessions. LongMemEval Any@3 measures answer-cluster retrieval across a larger haystack. Both are reported raw without cross-benchmark normalization.



