pollen-robotics/speech-commands-v0.02
收藏资源简介:
Speech Commands Dataset v0.02是一个用于有限词汇语音识别的数据集,包含一系列时长为一秒的.wav音频文件(16位PCM编码、16kHz采样率、单声道),每个文件包含一个单独的英语单词发音。总共有105,829个音频文件,按单词标签组织在文件夹中。数据集包含20个核心命令词:是、否、上、下、左、右、开、关、停、走、零、一、二、三、四、五、六、七、八、九;15个辅助词:床、鸟、猫、狗、快乐、房子、马文、希拉、树、哇、向后、向前、跟随、学习、视觉;以及一个背景噪声文件夹,包含环境噪声的较长音频片段。数据集按照原始存档中的官方validation_list.txt和testing_list.txt文件划分为训练集、验证集和测试集,确保同一发言者的所有话语被分配到同一分区。
A set of one-second .wav audio files (16-bit PCM, 16 kHz, mono), each containing a single spoken English word. 105,829 audio files organized into folders by word label. Words include 20 core command words: yes, no, up, down, left, right, on, off, stop, go, zero, one, two, three, four, five, six, seven, eight, nine; 15 auxiliary words: bed, bird, cat, dog, happy, house, marvin, sheila, tree, wow, backward, forward, follow, learn, visual; and background noise: the `_background_noise_` folder contains longer audio clips of environmental noise. The dataset is split into train/validation/test using the official `validation_list.txt` and `testing_list.txt` files included in the original archive, ensuring all utterances from a given speaker end up in the same partition.




