遇见数据集

Nexdata/Korean_Spontaneous_Speech_Data

收藏
Hugging Face2024-04-17 更新2024-03-04 收录
官方服务:

资源简介:

Nexdata/Korean_Spontaneous_Speech_Data数据集包含396小时的韩语自发语音数据,涵盖了多个主题。所有语音音频都经过人工转录为文本内容,并且标注了说话者身份、性别等信息。该数据集可用于声纹识别模型训练、机器翻译语料库构建以及算法研究。数据格式为16kHz、16bit、单声道,内容包括直播、综艺节目、演讲等。应用场景包括语音识别、视频字幕生成和视频内容审查。数据集的句子准确率(SAR)不低于95%。

The Nexdata/Korean_Spontaneous_Speech_Data dataset consists of 396 hours of Korean spontaneous speech data spanning diverse topics. All speech audio has been manually transcribed into text, with accompanying annotations including speaker identity, gender and other relevant information. This dataset can be applied to training speaker verification models, constructing machine translation corpora and conducting algorithmic research. The data is formatted as 16kHz, 16-bit, mono-channel, and covers content such as live broadcasts, variety shows, speeches and other materials. Its applicable scenarios include speech recognition, video subtitle generation and video content moderation. The Sentence Accuracy Rate (SAR) of the dataset is no less than 95%.

提供机构:
Nexdata
原始信息汇总

数据集卡片 Nexdata/Korean_Spontaneous_Speech_Data

描述

396小时 - 韩语即兴演讲数据集,内容涵盖多个主题。所有语音音频均已手动转录为文本内容;同时标注了说话者身份、性别等信息。该数据集可用于语音识别模型训练、机器翻译语料库构建及算法研究。

规格

格式

16kHz,16位,单声道;

内容类别

包括现场、综艺节目、演讲等;

语言

韩语;

标注

转录文本、说话者识别、性别标注;

应用场景

语音识别、视频字幕生成和视频内容审核;

准确性

句子准确率(SAR)不低于95%。

许可信息

商业许可

二维码
社区交流群
二维码
科研交流群
商业服务