ytc1997/multisource-membench
收藏资源简介:
Multi-Source Memory Benchmark(多源记忆基准)是一个诊断测试平台,用于在冲突的多源个人记忆上进行选择性问答(ANSWER/SKIP)。每个人物角色有五个证据流,这些证据流从单个潜在事件表投影而来,具有已知、受控的每源失真(包括偏差方向、丢失率和粒度)。这使得方法能够针对潜在真实情况(而非任何单一源)进行评估。该基准伴随论文《Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison》,并用于比较基线方法、结构化融合方法以及前沿大型语言模型(如GPT、Gemini、DeepSeek、Qwen3系列)。数据集完全合成,包含34,560个问题实例,覆盖5个主题(工作、饮食、社交、睡眠、锻炼)和8种推理类型。它旨在评估方法在冲突证据、缺失字段和主题依赖的自我报告偏差下的表现,并研究选择性问答中的跳过成本与错误成本之间的权衡。
Multi-Source Memory Benchmark is a diagnostic testbed for selective question-answering (ANSWER/SKIP) over conflicting multi-source personal memory. Each persona has five evidence streams projected from a single latent event table with known, controlled per-source distortions (bias direction, dropout rate, granularity), allowing methods to be measured against the latent ground truth rather than against any single source. The benchmark accompanies the paper Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison and is used to compare baselines, structured fusion methods, and frontier LLMs (GPT, Gemini, DeepSeek, Qwen3 families). The dataset is fully synthetic, with 34,560 question instances covering 5 topics (Work, Diet, Social, Sleep, Exercise) and 8 reasoning types. It is intended for evaluating methods under conflicting evidence, missing fields, and topic-dependent self-report bias, and for studying the cost-of-skip vs cost-of-wrong trade-off in selective QA.





