ales_razanau_output
收藏资源简介:
“Зборнік”数据集是Ministerskija集合的一部分,专门用于白俄罗斯语自动语音识别任务。该数据集由作者Алесь Разанаў创建,包含对齐的白俄罗斯语有声读物音频录音及其转录文本。数据集规模为1K到10K级别,具体包含106个数据行(其中93行在HuggingFace上发布),音频总时长约为21分钟。每个数据样本包含三个字段:audio(约15秒的音频片段)、text(对应的转录文本)和chunk_uid(片段的唯一标识符)。数据通过处理流程生成:原始有声读物被分割成短音频片段,然后使用Gemini和两个独立的自动语音识别系统与转录文本进行对齐,确保对齐置信度不低于0.95。该数据集适用于训练或评估白俄罗斯语语音识别模型,尤其针对有声读物领域的应用。
The "Зборнік" dataset is part of the Ministerskija Collection, specifically dedicated to Belarusian automatic speech recognition tasks. Created by author Ales Razanau, this dataset includes aligned Belarusian audiobook audio recordings and their corresponding transcriptions. The dataset has a scale ranging from 1K to 10K, containing 106 data entries in total, 93 of which are released on Hugging Face, with a total audio duration of approximately 21 minutes. Each data sample comprises three fields: `audio` (a ~15-second audio clip), `text` (the corresponding transcription text), and `chunk_uid` (the unique identifier of the clip). The dataset is generated through a processing pipeline: original audiobooks are split into short audio segments, then aligned with their transcriptions using Gemini and two independent automatic speech recognition systems, with an alignment confidence score of no less than 0.95. This dataset is applicable for training or evaluating Belarusian speech recognition models, especially for audiobook-domain applications.
数据集概述:Зборнік — Алесь Разанаў
- 语言:白俄罗斯语(Belarusian)
- 许可证:CC0-1.0(公有领域)
- 任务类别:自动语音识别(ASR)
- 标签:audiobook, belarusian, speech, asr, aligned, speaker_01
- 数据集大小:1,000条至10,000条之间
数据集详情
- 发布行数(HF):93行
- 数据库总计:106行
- 音频总时长:0小时21分钟
- 置信度阈值:≥ 0.95
数据结构
每一行包含三个字段:
audio:音频片段(约15秒)text:转录文本(由Gemini与ASR对齐生成)chunk_uid:片段的唯一标识符
数据处理
音频书籍被分割为短片段,并通过Gemini和两个独立的ASR系统进行转录对齐。对齐置信度≥0.95。
说话人信息
- 聚类:speaker_01
- 平均相似度评分:0.878
- 最相近数据集:raisa_baravikova_vasmiradkou_i_output(相似度0.93)
说话人识别由WavLM-Base+模型确定(余弦相似度,阈值=0.82)。
所属合集
该数据集属于 Ministerskija 收藏集,该收藏集包含已对齐的白俄罗斯语音频书转录数据。




