SLT 2026 SmartGlasses Challenge Dataset
收藏资源简介:
SLT 2026 SmartGlasses Challenge数据集是由西北工业大学等多所机构联合构建的面向智能眼镜的第一人称多说话人语音处理基准。该数据集包含714个真实场景会话,总时长106.98小时,涵盖双人对话和多人会议两大场景,采用四通道MEMS麦克风阵列录制。数据通过“大纲引导自发言语”协议采集,辅以LLM辅助生成对话大纲,确保自然性与语义复杂度,并经过人工标注与两轮交叉验证。数据集旨在评估带时间戳的说话人属性语音识别(TSA-ASR)和口语理解(SLU)任务,特别是在高重叠率、长上下文和动态声学环境下的表现,弥补了现有基准在可穿戴设备场景下的不足。
The SLT 2026 SmartGlasses Challenge Dataset is a first-person multi-speaker speech processing benchmark for smart glasses, jointly constructed by Northwestern Polytechnical University and several other institutions. It contains 714 real-world session recordings with a total duration of 106.98 hours, covering two primary scenarios: two-person dialogues and multi-person meetings, and was recorded using a four-channel MEMS microphone array. The data was collected via the outline-guided spontaneous speech protocol, where dialogue outlines were generated with the assistance of LLMs to guarantee naturalness and semantic complexity, and underwent manual annotation and two rounds of cross-validation. This dataset aims to evaluate tasks including timestamped speaker-attributed speech recognition (TSA-ASR) and spoken language understanding (SLU), particularly performance under high speech overlap, long context, and dynamic acoustic environments, thus filling the gap of existing benchmarks in wearable device application scenarios.





