遇见数据集

Run records for "Evidence Packing Under Token Budgets for Long-Term Conversational Question Answering"

收藏
Zenodo2026-09-26 更新2026-10-01 收录
官方服务:

资源简介:

Case-level run records for five studies of evidence packing in long-term conversational question answering: a three-arm packing ablation on the 500 questions of LongMemEval S (3,000 judged cases), an analysis of every answer that changed in its primary contrast, a date-order and context-window ablation (3,000 cases), a pre-specified held-out confirmation on the 1,540 answerable LoCoMo questions (9,240 cases), and a re-judging of all 15,240 cases with the official LongMemEval GPT-4o judge with blinded adjudication of disagreements. The package contains the frozen study protocols and amendments, per-case correctness labels, analysis summaries, raw-file SHA-256 manifests and cost receipts. It reproduces every count, interval and test statistic in the accompanying paper from the released records. It contains no code. LongMemEval (MIT) model answers are included. LoCoMo is licensed CC BY-NC 4.0 by its authors, so no text derived from LoCoMo conversations is included: LoCoMo records carry identifiers, labels and statistics only. Neither benchmark is redistributed; obtain them from their original sources. See README.md.

提供机构:
Zenodo
创建时间:
2026-09-26
二维码
社区交流群
二维码
科研交流群
商业服务