pasketti
收藏资源简介:
该数据集名为'On Top of Pasketti — Children's Word ASR (Training Data)',是一个用于儿童语音自动识别(ASR)任务的训练数据集。数据集包含95,572条语音样本,总时长约185.4小时,音频格式为FLAC。每条样本包含多个字段:utterance_id(唯一标识符)、child_id(匿名说话者标识符)、session_id(录音会话标识符)、audio_path(音频文件路径)、audio_duration_sec(音频时长)、age_bucket(年龄范围)、md5_hash(音频文件MD5校验和)、filesize_bytes(音频文件大小)、orthographic_text(规范化文本转录)和audio(嵌入的FLAC音频)。数据集还提供了年龄分布统计,其中8-11岁儿童的样本占比最高(77.4%)。该数据集适用于开发和评估针对儿童语音的ASR模型。
This dataset, named *On Top of Pasketti — Children's Word ASR (Training Data)*, is a training dataset for children's speech automatic speech recognition (ASR) tasks. It consists of 95,572 speech samples with a total duration of approximately 185.4 hours, and the audio is stored in FLAC format. Each sample contains the following fields: utterance_id (unique identifier), child_id (anonymous speaker identifier), session_id (recording session identifier), audio_path (path to the audio file), audio_duration_sec (audio duration in seconds), age_bucket (age group), md5_hash (MD5 checksum of the audio file), filesize_bytes (audio file size in bytes), orthographic_text (normalized text transcription), and audio (embedded FLAC audio). The dataset also provides age distribution statistics, where samples from children aged 8 to 11 account for the largest proportion (77.4%). This dataset is applicable to the development and evaluation of ASR models tailored for children's speech.



