VietSuperSpeech
收藏资源简介:
VietSuperSpeech 是一个越南语语音识别数据集,包含来自三个来源(asr_dataset_nguoivietdailynews、asr_dataset_nguyenkhangofficial、asr_dataset_trinhlieu)的语音数据。数据集总规模为32,267个样本,总时长103.18小时,采样率为16kHz,平均片段长度约12秒。数据划分为29,041个训练样本和3,226个开发样本。数据集采用Icefall格式组织,包含train.json(训练样本)、dev.json(开发样本)、manifest.json(元数据)和audio/目录(按来源组织的音频文件)。每个样本包含音频文件相对路径、转录文本、时长(秒)和来源视频/文件名信息。所有文本使用Zipformer-30M-RNNT-6000h模型进行转录。数据集采用MIT许可协议发布。
VietSuperSpeech is a Vietnamese automatic speech recognition (ASR) dataset containing speech data from three sources: asr_dataset_nguoivietdailynews, asr_dataset_nguyenkhangofficial, and asr_dataset_trinhlieu. The dataset has a total of 32,267 samples, with a total duration of 103.18 hours, a sampling rate of 16 kHz, and an average segment length of approximately 12 seconds. It is split into 29,041 training samples and 3,226 development samples. The dataset is organized in Icefall format, including train.json (for training samples), dev.json (for development samples), manifest.json (metadata), and an audio/ directory that organizes audio files by their sources. Each sample contains the relative path to the audio file, the transcribed text, the duration in seconds, and the source video or filename information. All transcriptions were generated using the Zipformer-30M-RNNT-6000h model. The dataset is released under the MIT License.



