该数据集名为'On Top of Pasketti — Children's Word ASR (Training Data)',是一个用于儿童语音自动识别(ASR)任务的训练数据集。数据集包含95,572条语音样本,总时长约185.4小时,音频格式为FLAC。每条样本包含多个字段:utterance_id(唯一标识符)、child_id(匿名说话者标识符)、session_id(录音会话标识符)
TC-STAR is a European integrated project focusing on Speech-to-Speech Translation (SST). To encourage significant breakthrough in all SST technologies, annual open competitive evaluations are organize