Vaani-English-preprocessed
收藏资源简介:
该数据集是一个多模态语音数据集,包含音频、文本转录和说话人/地域元数据。数据集包含15,918个样本,划分为训练集(12,729样本)、验证集(1,485样本)和测试集(1,704样本)。每个样本包含以下字段:音频数据(采样率16kHz)、语言标识、说话人性别、所属州/省、所属地区、原始转录文本、参考图像路径、未清理的无脱字符索引、清理后的转录文本、以及小写且无标点的清理转录文本。数据集适用于语音识别、语音合成、说话人属性分析、方言/地域语言研究等任务。数据总大小约2.22GB。
This is a multimodal speech dataset that includes audio, text transcripts, and speaker/regional metadata. The dataset comprises 15,918 total samples, divided into three subsets: a training set with 12,729 samples, a validation set with 1,485 samples, and a test set with 1,704 samples. Each sample contains the following fields: audio data (sampling rate: 16 kHz), language identifier, speaker gender, affiliated state or province, affiliated region, original transcription text, path to the reference image, uncleaned caret-free index, cleaned transcription text, and lowercased, punctuation-free cleaned transcription. This dataset is applicable to tasks such as speech recognition, speech synthesis, speaker attribute analysis, and dialect/regional language research. The total size of the dataset is approximately 2.22 GB.




