Komzeras
收藏资源简介:
该数据集是一个用于自动语音识别任务的布列塔尼语(Breton)语音数据集。数据来源于 Axel Landeau 在 YouTube 频道上发布的布列塔尼语视频录音,其对应的转录文本则发布在 Komzeras 博客上。数据集通过将原始音频与转录文本按 30 秒的音频块进行对齐处理而构建,总音频时长约为 1 小时 2 分钟 13 秒。数据集包含 128 个训练样本,每个样本由三个字段构成:title(标题,字符串类型)、audio(音频,采样率为 16000 Hz 的未解码音频数据)和 br(布列塔尼语转录文本,字符串类型)。数据集采用 CC-BY-SA-4.0 许可证发布。
This dataset is a Breton speech dataset for automatic speech recognition tasks. The data originates from Breton language video recordings published by Axel Landeau on a YouTube channel, with corresponding transcriptions released on the Komzeras blog. The dataset is constructed by aligning the original audio with transcription texts into 30-second audio chunks, with a total audio duration of approximately 1 hour, 2 minutes, and 13 seconds. It contains 128 training samples, each consisting of three fields: title (title, string type), audio (audio, undecoded audio data with a sampling rate of 16000 Hz), and br (Breton transcription text, string type). The dataset is released under the CC-BY-SA-4.0 license.
数据集概述
数据集名称: Komzeras
语言: 布列塔尼语(br)
任务类别: 自动语音识别(automatic-speech-recognition)
许可证: CC-BY-SA-4.0
数据集来源: 原始视频来自 YouTube 频道 Komzeras(作者 Axel Landeau),转录文本来自博客。
数据集构建方法: 将 YouTube 视频的录音与博客提供的转录文本进行对齐,按 30 秒音频块切分,最终形成此数据集。
数据集规模: 总时长 1 小时 2 分钟 13 秒
数据集特征列:
title: 字符串类型,视频标题。audio: 音频数据,采样率 16000 Hz,未解码(decode: false)。br: 字符串类型,布列塔尼语转录文本。
数据集划分:
- 训练集(train):包含 128 个样本,总字节数约 119.53 MB,下载大小约 119.47 MB。
数据文件路径: data/train-*(默认配置)





