ivan_ptashnikau_output
收藏资源简介:
该数据集名为Зборнік — Іван Пташнікаў,是Ministerskija收藏的一部分,专门用于白俄罗斯语自动语音识别任务。它是一个经过严格对齐的白俄罗斯语有声读物音频及其文本转录数据集,作者为Іван Пташнікаў。数据集包含总计1,068条记录,其中在HuggingFace平台上公开发布了892条。音频内容总时长约为2小时8分钟。每条数据记录由三个核心字段构成:audio字段存储时长约15秒的音频片段;text字段存储对应的白俄罗斯语文本转录;chunk_uid字段是每个片段的唯一标识符。数据集的构建经过了精细的处理流程:原始有声读物音频被分割成短片段,然后利用Gemini工具结合两个独立的自动语音识别系统进行音频与文本的对齐处理,并确保所有收录数据的对齐置信度不低于0.95。该数据集适用于训练或评估白俄罗斯语的自动语音识别模型,尤其适用于有声读物领域的语音研究。
The dataset is named Зборнік — Іван Пташнікаў and is part of the Ministerskija collection, focusing on Belarusian automatic speech recognition tasks. It is a rigorously aligned dataset of Belarusian audiobook audio and its text transcriptions, authored by Іван Пташнікаў. The dataset contains a total of 1,068 records, with 892 publicly released on the HuggingFace platform. The total audio duration is approximately 2 hours and 8 minutes. Each data record consists of three core fields: the audio field stores audio segments of about 15 seconds in duration; the text field stores the corresponding Belarusian text transcription; and the chunk_uid field is a unique identifier for each segment. The dataset construction involved a meticulous processing pipeline: the original audiobook audio was split into short segments, then aligned with text using the Gemini tool combined with two independent automatic speech recognition systems, ensuring an alignment confidence of at least 0.95 for all included data. This dataset is suitable for training or evaluating Belarusian automatic speech recognition models, particularly for speech research in the audiobook domain.
数据集概述
- 数据集名称: Зборнік — Іван Пташнікаў
- 语言: 白俄罗斯语 (be)
- 许可证: CC0-1.0
- 任务类别: 自动语音识别 (automatic-speech-recognition)
- 标签: audiobook, belarusian, speech, asr, aligned, speaker_01
- 大小类别: 1K < n < 10K
数据集规模
| 指标 | 数值 |
|---|---|
| 已发布行数 (HF) | 892 |
| 数据库总数 | 1,068 |
| 音频总时长 | 2小时08分钟 |
| 置信度阈值 | ≥ 0.95 |
数据结构
每一行包含:
audio— 音频片段(约15秒)text— 转录文本(由 Gemini 和 ASR 对齐生成)chunk_uid— 片段的唯一标识符
数据处理
音频书被分割成短片段,并通过 Gemini 和两个独立的 ASR 系统与转录文本对齐。对齐置信度 ≥ 0.95。
说话人信息
| 指标 | 数值 |
|---|---|
| 聚类 | speaker_01 |
| 平均相似度评分 | 0.923 |
| 最接近的数据集 | stefan_tsvei_g_nyabachnaya_kalektsyya_output (0.96) |
说话人由 WavLM-Base+ 模型(余弦相似度,阈值=0.82)识别。
所属合集
该数据集是 Ministerskija 合集的一部分,该合集包含对齐的白俄罗斯语音频书录音与转录文本。




