遇见数据集

Per-question outcomes for M3 Memory on LongMemEval-M (n=500)

收藏
Zenodo2026-05-12 更新2026-05-29 收录
官方服务:

资源简介:

Per-question outcome vectors from a full-corpus evaluation of M3 Memory (an open-source SQLite-backed conversation memory system) on LongMemEval-M (Wu et al., ICLR 2025; n=500 questions x ~1.5 M tokens per problem). This dataset is the master per-question artifact underlying every cell in section 4.1, 4.2, 4.5, 4.6, 4.9, 5, and 4.10 of the accompanying manuscript ("M3 Memory on LongMemEval-M: Retrieval and Augmentation at 1.5 Million Token Scale"). The CSV has 500 rows x 28 columns: identity fields (question_id, qtype, is_abstention, n_gold_sessions); gold-in-context flags (turn-level and session-level on raw-tier top-20); end-to-end QA correctness across 5 reader configurations x 2 prompt variants x 3 judge families; parametric floor and reader-ceiling-under-oracle-retrieval correctness; SHR@20 indicators across 5 retrieval configurations. All 22 boolean columns reproduce the corresponding manuscript cell to plus-or-minus 0.001; see per_question_outcomes.validation.md for the per-column reproducibility table. The underlying LongMemEval-M corpus is NOT redistributed here; this dataset references only per-question identifiers from the canonical Wu et al. release at github.com/xiaowu0162/LongMemEval. Reviewers and replicators can re-bootstrap any cell in the paper from this single artifact using the recipe in per_question_outcomes.README.md (paired bootstrap, 10,000 reps, seed 42; matches bin/paper_pvalues.py in the accompanying repository at github.com/skynetcmd/m3-memory).

提供机构:
Zenodo
创建时间:
2026-05-12
二维码
社区交流群
二维码
科研交流群
商业服务