Finnish Parliament ASR corpus
收藏资源简介:
Finnish Parliament ASR corpus是由阿尔托大学创建的,目前最大的公开可用芬兰语自动语音识别(ASR)数据集,包含超过3000小时的手动转录语音数据和449位发言者的丰富人口统计元数据。该数据集基于早期的初步工作,自然地分为两个训练子集,分别来自两个不同的时间段,并提供了两个官方的修正测试集,覆盖不同的时间段,设置了具有纵向分布变化特性的ASR任务。此外,还提供了一个官方开发集。数据集的应用领域包括ASR系统的训练和评估,以及解决语音识别中的性别、年龄和教育水平偏差问题。
The Finnish Parliament ASR Corpus, created by Aalto University, is currently the largest publicly available Finnish automatic speech recognition (ASR) dataset. It contains over 3,000 hours of manually transcribed speech data and rich demographic metadata for 449 speakers. Built upon early preliminary work, the corpus is naturally split into two training subsets sourced from two distinct time periods, and provides two official revised test sets covering different timeframes, establishing an ASR task with longitudinal distribution variations. Furthermore, an official development set is provided. The dataset is applied to the training and evaluation of ASR systems, as well as the mitigation of gender, age, and education level biases in speech recognition.

- 1Finnish Parliament ASR corpus - Analysis, benchmarks and statistics阿尔托大学 · 2022年



