raisa_baravikova_vershy_pra_kahanne_output
收藏资源简介:
该数据集名为“Вершы пра каханне — Раіса Баравікова”,是白俄罗斯诗人Raisa Baravikova爱情诗的有声读物集合,属于Ministerskija系列的一部分,专门提供对齐的白俄罗斯语有声读物音频及其转录文本,用于自动语音识别(ASR)任务。数据集包含150个音频-文本对(其中122个已发布在Hugging Face平台),总音频时长约为31分钟。每个数据样本包括一个约15秒的音频片段(audio)、对应的转录文本(text)以及一个唯一的片段标识符(chunk_uid)。数据通过将原始有声读物分割成短片段,并利用Gemini模型结合两个独立的ASR系统进行音频与文本的对齐处理生成,置信度阈值设定为不低于0.95,确保了高质量的对齐结果。该数据集适用于白俄罗斯语的语音识别模型训练、评估以及相关语音技术研究。
The dataset is named Вершы пра каханне — Раіса Баравікова and is a collection of audiobooks featuring love poems by Belarusian poet Raisa Baravikova. It is part of the Ministerskija series, dedicated to providing aligned Belarusian audiobook audio and transcriptions specifically for automatic speech recognition (ASR) tasks. The dataset contains 150 audio-text pairs (with 122 released on the Hugging Face platform), with a total audio duration of approximately 31 minutes. Each data sample consists of an audio segment of about 15 seconds (audio), the corresponding transcription text (text), and a unique segment identifier (chunk_uid). The data was generated by splitting the original audiobook into short segments and aligning audio and text using the Gemini model combined with two independent ASR systems, ensuring high-quality alignment with a confidence threshold of at least 0.95. This dataset is suitable for training and evaluating Belarusian speech recognition models, as well as for related speech technology research.
-
语言: 白俄罗斯语 (be)
-
许可协议: CC0-1.0 (cc0-1.0)
-
任务类别: 自动语音识别 (automatic-speech-recognition)
-
标签: 有声书、白俄罗斯语、语音、ASR、对齐、speaker_01
-
数据集规模: 1,000 到 10,000 条 (1K<n<10K)
-
正式名称: Вершы пра каханне — Раіса Баравікова
-
作者: Раіса Баравікова
-
所属合集: Ministerskija — 白俄罗斯语有声书的校准音频与转录本集合
-
数据统计:
- 已发布的行数 (HF): 122
- 数据库总数: 150
- 音频总时长: 0小时31分钟
- 置信度阈值: ≥ 0.95
-
数据结构:
- 每条数据包含:
audio: 音频片段 (约15秒)text: 转录文本 (由Gemini + ASR对齐生成)chunk_uid: 片段的唯一标识符
- 每条数据包含:
-
数据处理:
- 有声书被分割成短片段,并使用Gemini和两个独立的ASR系统与转录文本进行对齐。
- 对齐置信度 ≥ 0.95。
-
说话人信息:
- 说话人聚类: speaker_01
- 平均相似度评分: 0.899
- 最接近的数据集: genadz_pashkou_output (相似度 0.95)
- 说话人识别模型: WavLM-Base+ (使用余弦相似度,阈值为0.82)




