kyutai/interactivity-alignment-samples
收藏资源简介:
该数据集包含论文《多层面交互对齐在全双工语音模型中》的音频样本。这些样本是在Full-Duplex-Bench v1(使用预录制输入的静态评估)和Full-Duplex-Bench v2(与GPT-Realtime进行实时多轮对话)上生成的,用于博客文章和演示页面。每个文件都是立体声WAV格式,其中通道0承载输入说话者,通道1承载模型的输出。样本比较了基础模型(Moshi、PersonaPlex)与经过强化学习后训练版本(在Fisher和Seamless Interaction数据集上训练)的表现。数据集按任务和变体组织,包括合成暂停处理、坦诚轮转、ICC反馈、合成用户中断等任务,以及不同模型变体的音频文件。
This repository hosts the audio samples generated on Full-Duplex-Bench v1 (static evaluation with pre-recorded input) and Full-Duplex-Bench v2 (real-time multi-turn dialogue with GPT-Realtime), used in the blog post and demo page. Each file is a stereo WAV where channel 0 carries the input speaker and channel 1 carries the models output. The samples compare base models (Moshi, PersonaPlex) against our RL post-trained versions on Fisher and Seamless Interaction datasets. The dataset is organized by tasks and variants, including tasks such as synthetic pause handling, candor turn taking, ICC backchannel, synthetic user interruption, and audio files for different model variants.




