BSC-LT/distilled-catalan-youtube-speech
收藏资源简介:
Distilled Catalan YouTube Speech Corpus是一个精选的加泰罗尼亚语YouTube语音语料库子集,包含207小时的转录语音。该数据集通过两个独立的自动语音识别(ASR)系统(验证模型)自动验证转录内容,并根据系统一致性对转录质量进行分类:完美匹配(系统输出完全相同)和单词计数匹配(单词数相同但措辞不同,通过第三个ASR系统解决)。此外,还提供了手动标注的测试集和高置信度的验证集。该数据集旨在支持自动语音识别(ASR)的研究和开发,特别是在低资源和半监督设置下。数据集包含音频ID、音频文件、语料库ID、分割信息、语言、持续时间、性别、YouTube链接、共识信息、选择转录和规范化文本等字段。数据集分为perfect_matches、word_count_matches、validation和test四个部分。数据集由巴塞罗那超级计算中心策划,采用MIT许可证。
The Distilled Catalan YouTube Speech Corpus is a curated subset of the Catalan YouTube Speech Corpus, containing 207 hours of transcribed speech. The dataset uses two independent automatic speech recognition (ASR) systems (verification models) to automatically verify transcriptions and categorizes transcription quality based on system agreement: perfect matches (identical outputs between systems) and word count matches (same word count but different wording, resolved using a third ASR system). Additionally, a manually annotated test set and a high-confidence validation set are provided. The dataset is intended to support research and development in automatic speech recognition (ASR), especially in low-resource and semi-supervised settings. The dataset includes fields such as audio ID, audio file, corpus ID, split information, language, duration, gender, YouTube URL, consensus information, selected transcription, and normalized text. The dataset is divided into four parts: perfect_matches, word_count_matches, validation, and test. The dataset was curated by the Barcelona Supercomputing Center and is licensed under MIT.




