VoxMem
收藏资源简介:
VoxMem是一个用于评估大型音频语言模型多模态记忆的基准数据集,涵盖语音语义、说话人身份、副语言线索和环境声音四种声学证据类型,以及信息提取、多会话推理、时间演化追踪和答案拒绝四种记忆操作,包含多会话口述历史,上下文长度从8K到64K tokens。
VoxMem is a benchmark dataset for evaluating multimodal memory of large audio-language models. It covers four types of acoustic evidence: speech semantics, speaker identity, paralinguistic cues, and environmental sounds, as well as four memory operation tasks: information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal. The dataset includes multi-session oral histories, with context lengths ranging from 8K to 64K tokens.
VoxMem 数据集概述
基本信息
- 名称:VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
- 作者:Yang Xiao、Vidhyasaharan Sethu、Eun-Jung Holden、Ting Dang
- 机构:University of Melbourne、University of New South Wales、AIMS Lab
- 论文:arXiv:2609.32607
- 发布时间:2026-09
- 数据集地址:https://huggingface.co/datasets/AudioMemory/voxmembench
- 项目主页:https://swagshaw.github.io/voxmem/
- 代码许可:MIT
- 数据许可:CC BY-NC 4.0
核心目标
评估大音频语言模型(LALMs)对多会话语音历史的记忆能力,不仅关注模型是否记住“说了什么”,还关注是否记住“谁说的”、“怎么说的”以及“听到了什么”——即仅存在于音频信号中、无法从文本转录中恢复的信息。
数据集规模
| 指标 | 数值 |
|---|---|
| 评估实例 | 3,196(799 问题 × 4 种上下文长度) |
| 语音会话 | 34,743(177 小时音频) |
| 上下文预算 | 8K、16K、32K、64K tokens |
| 有效分类单元 | 15 个(证据类型 × 记忆操作) |
| 评估模型 | 15 个 LALMs(5 个专有模型、10 个开源模型) |
评估框架
两个评估维度
声学证据类型(Acoustic evidence)——需要恢复什么:
| 类型 | 含义 |
|---|---|
speech_semantics |
用户说了什么 |
speaker_information |
谁在说话 |
paralinguistic_information |
怎么说的——声音表现 |
environmental_sound |
用户周围能听到什么 |
记忆操作(Memory operation)——如何使用:
| 操作 | 含义 |
|---|---|
information_extraction (IE) |
从单个会话中恢复一个事实 |
multi_session_reasoning (MSR) |
跨会话整合证据 |
temporal_evolution_tracking (TET) |
追踪某事物随时间的变化 |
answer_refusal (AR) |
识别历史记录中不包含答案 |
两个维度交叉得到 15 个有效单元,其中 speech_semantics × information_extraction 被有意留空。
排行榜(32K 参考预算,总体准确率 %)
| 排名 | 模型 | 类型 | 语义 | 说话人 | 副语言 | 环境声 | 总体 |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8-Omni-Flash | 专有 | 60.3 | 42.2 | 25.4 | 25.0 | 38.5 |
| 2 | Gemini-3.1-Pro | 专有 | 68.5 | 29.3 | 19.7 | 21.2 | 35.9 |
| 3 | Gemini-3.8-Flash | 专有 | 54.3 | 41.5 | 20.8 | 19.2 | 34.0 |
| 4 | Qwen3.5-Omni-Plus | 专有 | 56.0 | 31.3 | 19.3 | 23.7 | 33.0 |
| 5 | Qwen3-Omni-30B-A3B | 开源 | 38.4 | 25.2 | 20.1 | 18.6 | 26.0 |
| 6 | Ultravox-v0.6-Llama-3.1-8B | 开源 | 35.3 | 31.3 | 15.5 | 17.3 | 24.5 |
| 7 | Gemini-2.5-Pro | 专有 | 38.8 | 19.0 | 14.8 | 20.5 | 23.7 |
| 8 | Gemma-4-E4B-it | 开源 | 35.8 | 29.3 | 14.0 | 16.7 | 23.7 |
| 9 | Baichuan-Audio-7B | 开源 | 30.6 | 35.4 | 13.3 | 13.5 | 22.4 |
| 10 | MiniCPM-o-4.5 | 开源 | 31.0 | 28.6 | 11.4 | 18.6 | 21.7 |
| 11 | Audio-Flamingo-Next | 开源 | 31.9 | 23.8 | 15.2 | 13.5 | 21.3 |
| 12 | FireRedAudio-9B | 开源 | 28.9 | 25.9 | 15.2 | 14.7 | 21.0 |
| 13 | Phi-4-Multimodal | 开源 | 27.6 | 27.9 | 14.0 | 13.5 | 20.4 |
| 14 | MiMo-Audio-7B-Instruct | 开源 | 27.2 | 23.1 | 13.3 | 18.6 | 20.2 |
| 15 | Baichuan-Omni-1.5-7B | 开源 | 28.9 | 17.0 | 12.9 | 10.3 | 17.8 |
| 专有模型均值(5) | 55.6 | 32.7 | 20.0 | 21.9 | 33.0 | ||
| 开源模型均值(10) | 31.6 | 26.7 | 14.5 | 15.5 | 21.9 |
关键发现
- 当前 LALMs 远未实现可靠的语音对话记忆。32K 下无模型总体超过 40%,最优仅 38.5%。
- 非词汇声学信息的可及性远低于语音语义。专有模型在语义上平均 55.6%,但在说话人、副语言、环境证据上仅为 32.7%、20.0%、21.9%。
- 难度同时取决于操作与证据类型。时序追踪对语义有效(44.5%),但对副语言(3.4%)和环境(1.2%)证据近乎崩溃。
- 答案拒绝与作答表现出不同的模式:模型恰好在难以利用声学证据之处更容易成功弃答。
- 随着历史增长,对同一证据的获取能力下降,且不同证据类型下降速率不同。
- 错误特征存在质的差异。说话人错误多为绑定失败(48%);副语言错误多为定位失败(63%)。
数据结构与配置
每个数据行即一个基准项目,包含运行所需的全部内容——有序的会话历史、内联用户音频片段以及口述问题。数据集提供以下配置:
| 配置 | 条目数 | 8K | 16K | 32K | 64K |
|---|---|---|---|---|---|
<length> |
799 | 4.03 GB | 7.72 GB | 15.38 GB | 32.83 GB |
<length>_speech_semantics |
232 | 1.19 GB | 2.26 GB | 4.50 GB | 9.18 GB |
<length>_speaker_information |
147 | 0.73 GB | 1.41 GB | 2.80 GB | 6.52 GB |
<length>_paralinguistic_information |
264 | 1.33 GB | 2.55 GB | 5.12 GB | 10.38 GB |
<length>_environmental_sound |
156 | 0.79 GB | 1.50 GB | 2.96 GB | 6.75 GB |
<length> 为 8k、16k、32k 或 64k。
提示词构成
模型输入依次为:系统提示词、按顺序排列的历史记录、口述问题。每个会话以用户轮次开头,先以文本形式给出 Session timestamp: ...,随后是该轮次的音频;后续用户轮次仅含音频,助手轮次为文本。运行器绝不会将转录文本放入提示词中。
评分方式
- 指标:准确率,分为两个永不合并平均的层:
answerable(2,676 项):当响应在答案归一化下与gold_json含义相同时判为正确。answer_refusal(520 项):当响应表示证据不足时判为正确。
- 使用 LLM 判官判定等价性,判官仅看到问题、金标准与响应,不接触音频与历史。
- 报告还按上下文长度和 15 个证据类型 × 记忆操作单元细分准确率。
输出格式
predictions.jsonl,每项一个对象,包含评分所需的全部信息,示例如下:
json {"item_id": "q_00006_8k", "question_id": "q_00006", "context_length": "8K", "evidence_type": "speaker_information", "memory_operation": "temporal_evolution_tracking", "expected_response": "answer", "gold_json": "[...]", "answer_type": "ordered_list", "response": "...", "n_audio_clips": 51, "verdict": true, "judge": "gpt-4o-mini"}
基准设计要点
- 上下文长度:每个问题在 8K、16K、32K、64K 四个音频 token 长度下出现,共享同一
question_id,问题、金标准和证据完全一致,仅周围历史增长,且历史嵌套(sessions(8K) ⊆ sessions(16K) ⊆ ...)。 - 干扰项:每个历史按会话标注——
evidence(证据会话)、samekey_haystack(共享检索键、仅靠问题指定的选择器排除)、topical_haystack(共享主题但不共享查询属性)、filler。
引用
bibtex @article{xiao2026voxmem, title = {VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models}, author = {Xiao, Yang and Sethu, Vidhyasaharan and Holden, Eun-Jung and Dang, Ting}, journal = {arXiv preprint arXiv:2609.32607}, year = {2026} }





