iagan_frydryh_shyler_kubak_output
收藏资源简介:
“Кубак — Іаган Фрыдрых Шылер”数据集是Ministerskija收藏的一部分,专门用于白俄罗斯语的自动语音识别任务。该数据集包含白俄罗斯语有声读物(内容源自作者约翰·弗里德里希·席勒)的音频片段及其对应的文本转录,并经过严格的对齐处理。数据集规模包含104条记录(在HuggingFace上发布102条),音频总时长约为14分钟。每条数据记录由三个字段构成:audio字段为时长约15秒的音频片段;text字段为对应的白俄罗斯语文本转录;chunk_uid字段为片段的唯一标识符。数据通过将原始有声读物切分为短片段,并利用Gemini模型结合两套独立的自动语音识别系统进行文本对齐而生成,确保对齐置信度不低于0.95。该数据集适用于训练或评估白俄罗斯语的自动语音识别模型,尤其适用于有声读物场景下的语音转文本任务。
The "Kubak — Johann Friedrich Schiller" dataset is part of the Ministerskija Collection, specifically designed for Belarusian automatic speech recognition (ASR) tasks. It consists of audio clips and their corresponding transcriptions of Belarusian audiobooks (content derived from the works of Johann Friedrich Schiller), which have undergone strict alignment processing. The dataset contains 104 total records, with 102 released on HuggingFace, and the total audio duration is approximately 14 minutes. Each data record includes three fields: the `audio` field is an approximately 15-second audio clip; the `text` field is the matching Belarusian text transcription; and the `chunk_uid` field is the unique identifier of the clip. The dataset is generated by splitting the original audiobooks into short segments, and performing text alignment using the Gemini model combined with two independent automatic speech recognition systems, ensuring an alignment confidence of no less than 0.95. This dataset is applicable for training or evaluating Belarusian automatic speech recognition models, particularly for speech-to-text tasks in audiobook scenarios.
数据集详情:Кубак — Іаган Фрыдрых Шылер
基本信息
- 数据集名称: Кубак
- 语言: 白俄罗斯语 (Belarusian,
be) - 许可证: CC0-1.0
- 任务类型: 自动语音识别 (ASR)
- 标签: 有声书、白俄罗斯语、语音、ASR、对齐、speaker_03
- 数据规模: 1,000 < 样本数 < 10,000
- 发布行数 (HF): 102 行
- 数据库总行数: 104 行
- 音频总时长: 14分钟
- 对齐置信度阈值: ≥ 0.95
数据来源
该数据集是 Ministerskija 收藏集的一部分,包含已对齐的白俄罗斯语有声书音频片段及其转录文本。
数据结构
每一行数据包含以下字段:
audio— 音频片段(约15秒)text— 转录文本(使用 Gemini 和 ASR 对齐生成)chunk_uid— 片段唯一标识符
数据处理
有声书被分割为短片段,并使用 Gemini 和两个独立的 ASR 系统进行对齐对齐置信度 ≥ 0.95。
说话人信息
| 属性 | 值 |
|---|---|
| 聚类类别 | speaker_03 |
| 平均相似度评分 | 0.89 |
| 最接近的数据集 | astryd_lindgren_braty_lvinae_sertsa_output (相似度 0.97) |
说话人识别基于 WavLM-Base+ 模型(余弦相似度,阈值 0.82)。




