ernest_heminguei_stary_chalavek_i_mora_output
收藏资源简介:
该数据集名为Стары чалавек і мора(《老人与海》),是欧内斯特·海明威经典小说的白俄罗斯语译本有声读物。它属于Ministerskija集合,专门提供经过对齐处理的白俄罗斯语有声读物音频及其对应转录文本。数据集将完整有声读物分割成约15秒时长的音频片段,每个片段都配有通过Gemini和两个独立自动语音识别(ASR)系统进行高置信度(≥0.95)对齐后生成的准确转录文本。数据集共包含1,274条样本(其中在HuggingFace平台公开发布995条),音频总时长约为4小时30分钟。每个数据样本由三个字段构成:audio字段存储音频片段,text字段存储对应的转录文本,chunk_uid字段则是该片段的唯一标识符。该数据集主要用于白俄罗斯语的自动语音识别任务,为模型训练和评估提供了高质量的音频-文本对齐数据。
The dataset is named Стары чалавек і мора (The Old Man and the Sea), which is an audiobook of the Belarusian translation of Ernest Hemingways classic novel. It belongs to the Ministerskija collection, specifically providing aligned Belarusian audiobook audio and corresponding transcriptions. The dataset contains audio segments of approximately 15 seconds each, split from the complete audiobook, with each segment accompanied by accurate transcriptions generated through high-confidence (≥0.95) alignment using Gemini and two independent automatic speech recognition (ASR) systems. It includes a total of 1,274 samples (with 995 publicly released on HuggingFace), with an overall audio duration of about 4 hours and 30 minutes. Each data sample consists of three fields: audio for storing the audio segment, text for the corresponding transcription, and chunk_uid as the unique identifier for the segment. This dataset is primarily used for Belarusian automatic speech recognition tasks, providing high-quality audio-text aligned data for model training and evaluation.
数据集概述:Стары чалавек і мора
- 数据集名称:Стары чалавек і мора (The Old Man and the Sea)
- 作者:Эрнэст Хемінгуэй (Ernest Hemingway)
- 语言:白俄罗斯语 (Belarusian)
- 许可证:CC0-1.0(公共领域)
- 任务类别:自动语音识别 (Automatic Speech Recognition, ASR)
- 标签:audiobook, belarusian, speech, asr, aligned, speaker_02
- 数据规模:1,000 到 10,000 条样本 (1K<n<10K)
- 数据集大小:
- 已发布行数 (Hugging Face):995
- 数据库总行数:1,274
- 音频总时长:4小时30分钟
- 对齐置信度阈值:≥ 0.95
数据结构
每行数据包含三个字段:
- audio:音频片段,时长约15秒
- text:转录文本(通过 Gemini 和 ASR 对齐生成)
- chunk_uid:片段唯一标识符
数据处理说明
该数据集来自 Ministerskija 集合,是白俄罗斯有声书的对齐音频和转录文本。有声书被分割成短片段,并利用 Gemini 和两个独立的 ASR 系统进行转录对齐,对齐置信度不低于 0.95。
朗读者 (Speaker) 信息
- 说话人聚类:speaker_02
- 平均相似度评分:0.933
- 最相似数据集:kuzma_chorny_poshuki_buduchyni_output(相似度 0.97)
说话人身份通过 WavLM-Base+ 模型(余弦相似度,阈值=0.82)确定。




