ales_zhuk_praklytaya_lyubow_output
收藏资源简介:
该数据集是白俄罗斯语自动语音识别数据集,属于Ministerskija收藏的一部分,包含对齐的白俄罗斯语音频书籍录音及其转录文本。数据集基于白俄罗斯作家Алесь Жук的音频书籍《Пракляты любоў》创建。音频书籍被分割成约15秒的短片段,并使用Gemini和两个独立的自动语音识别系统与转录文本进行对齐处理,对齐置信度阈值设定为≥0.95。数据集包含1,038个音频片段,其中854个在HuggingFace平台上发布,音频总时长为3小时33分钟。每个数据样本包含三个字段:audio(音频片段)、text(对应的转录文本)和chunk_uid(片段的唯一标识符)。该数据集适用于白俄罗斯语自动语音识别任务的研究与开发,特别适用于音频书籍场景下的语音识别模型训练和评估。
This is a Belarusian automatic speech recognition (ASR) dataset, part of the Ministerskija collection, which contains aligned Belarusian audiobook recordings and their corresponding transcriptions. The dataset is developed based on the audiobook *Пракляты любоў* by Belarusian writer Алесь Жук. The original audiobook was segmented into short clips of approximately 15 seconds, and aligned with their transcriptions using Gemini and two independent ASR systems, with an alignment confidence threshold set to ≥0.95. The dataset includes 1,038 audio segments, 854 of which are released on the HuggingFace platform, with a total audio duration of 3 hours and 33 minutes. Each data sample contains three fields: `audio` (the audio segment), `text` (the corresponding transcription text), and `chunk_uid` (the unique identifier of the segment). This dataset is suitable for research and development of Belarusian automatic speech recognition tasks, and is particularly applicable to the training and evaluation of speech recognition models in audiobook scenarios.
数据集概述
数据集名称: Пракляты любоў (Cursed Love) — Алесь Жук (Ales Zhuk)
语言: 白俄罗斯语 (Belarusian)
许可证: CC0-1.0
任务类别: 自动语音识别 (Automatic Speech Recognition)
标签: 有声书, 白俄罗斯语, 语音, ASR, 对齐, speaker_02
数据规模: 1,000 < 样本数 < 10,000
数据集详情
- 作者: Алесь Жук (Ales Zhuk)
- 所属合集: Ministerskija — 白俄罗斯语有声书的对齐音频录音及转录集合。
- 已发布行数 (HF): 854 条
- 数据库总行数: 1,038 条
- 音频时长: 3小时33分钟
- 置信度阈值: ≥ 0.95
数据结构
每条记录包含以下字段:
audio— 音频片段(约15秒)text— 转录文本(通过 Gemini 和 ASR 对齐)chunk_uid— 片段的唯一标识符
数据处理
- 有声书被分割成短片段,并利用 Gemini 和两个独立的 ASR 系统与转录文本进行对齐。
- 对齐置信度 ≥ 0.95。
说话人信息
- 说话人聚类: speaker_02
- 平均相似度评分: 0.949
- 最接近的数据集: kuzma_chorny_zyamlya_output(相似度 0.97)
- 识别模型: WavLM-Base+(余弦相似度,阈值=0.82)




