SQA和CSI
收藏资源简介:
SQA(Speech Question Answering)数据集和CSI(Customized Speech Instruction)数据集是专为训练VLAS模型而创建的。SQA数据集由23K个场景组成,每个场景包含389个文本指令和194K个音频指令,覆盖了500种不同的声音,用于训练模型理解和执行基于语音的指令。CSI数据集则包含了500个不同声音的语音指令,用于增强模型在个性化语音指令理解方面的能力。这两个数据集的构建旨在推动机器人操作中语音指令的端到端处理技术的发展。
SQA (Speech Question Answering) and CSI (Customized Speech Instruction) datasets are specially developed for training the VLAS model. The SQA dataset comprises 23K scenarios, each containing 389 text instructions and 194K audio instructions, covering 500 distinct sound types, and is designed to train models to comprehend and execute speech-based instructions. The CSI dataset, by contrast, includes speech instructions derived from 500 different sounds, aiming to enhance the model’s capability in understanding personalized speech instructions. The construction of these two datasets is intended to advance the development of end-to-end processing technologies for speech instructions in robotic manipulation.




