遇见数据集

BEA-Base

收藏
arXiv2022-02-02 更新2024-06-21 收录
数据链接:
官方服务:

资源简介:

BEA-Base数据集由匈牙利语言研究中心创建,专注于自发匈牙利语音的自动语音识别(ASR)评估。该数据集包含140名不同背景的发言者的自发语音,旨在为对话AI应用提供评估基准。数据集的创建过程涉及从原始BEA数据库中精选数据,并进行必要的预处理以适应ASR实验。BEA-Base的应用领域主要集中在提高匈牙利语音识别系统的训练和评估,特别是在自发语音处理方面,以解决现有数据集在自发语音评估方面的不足。

The BEA-Base dataset, developed by the Research Centre for Linguistics of Hungary, focuses on automatic speech recognition (ASR) evaluation for spontaneous Hungarian speech. It contains spontaneous speech data from 140 speakers with diverse backgrounds, and aims to serve as an evaluation benchmark for conversational AI applications. The construction of BEA-Base involves selecting subsets from the original BEA database and conducting necessary preprocessing to adapt to ASR experimental requirements. The primary application scope of BEA-Base is centered on improving the training and evaluation of Hungarian speech recognition systems, especially for spontaneous speech processing, to address the deficiencies of existing datasets in spontaneous speech evaluation.

创建时间:
2022-02-02
二维码
社区交流群
二维码
科研交流群
商业服务