Seamless Interaction dataset
收藏资源简介:
SPEARBench基准构建于Seamless Interaction数据集之上,该数据集由研究团队从人类对话语料库中精心构建,旨在评估流式语音到语音语言模型的自然度。数据集包含5419条精选的双人对话片段,总计约37.33小时音频,源自2007段原始对话,覆盖即兴和自然主义两种子集,数据来源于真实的人类互动录音,并经过转录和说话人分离处理。其创建过程通过特定协议从原始对话中提取包含上下文、问题和回答的有效对话结构,确保每个片段均以问题结束并配有真实人类答案。该数据集主要应用于语音到语音模型的自然度评估领域,旨在系统量化模型在响应延迟、语音质量、方言一致性、情感自然度及人际姿态等多维度的表现,以解决现有基准在全面衡量对话行为自然性方面的不足。
SPEARBench is built upon the Seamless Interaction dataset, which was meticulously constructed by a research team from human conversational corpora to evaluate the naturalness of streaming speech-to-speech large language models. The dataset includes 5,419 curated two-party dialogue segments, totaling approximately 37.33 hours of audio, derived from 2,007 original conversations and covering two subsets: impromptu and naturalistic dialogues. The data originates from real human interaction recordings, and has undergone transcription and speaker diarization processing. Its creation process extracts valid dialogue structures containing context, questions and answers from original conversations via a specific protocol, ensuring that each segment ends with a question and is paired with a genuine human response. This dataset is primarily applied in the field of naturalness evaluation for speech-to-speech models, aiming to systematically quantify model performance across multiple dimensions including response latency, speech quality, dialect consistency, emotional naturalness and interpersonal stance, so as to address the shortcomings of existing benchmarks in comprehensively measuring the naturalness of conversational behaviors.




