遇见数据集

jimregan/clarinpl_sejmsenat

收藏
Hugging Face2023-01-22 更新2024-03-04 收录
官方服务:

资源简介:

ClarinPL Sejm/Senat Speech Corpus数据集包含了97小时的波兰议会演讲录音,这些录音发布在ClarinPL网站上。数据集中的每个数据点通常包括音频文件的路径(通常称为`file`)及其转录文本(称为`text`)。音频文件为波兰语。数据集的结构包括数据实例、数据字段和数据分割。数据实例示例展示了音频文件路径、ID、说话者ID和转录文本。数据字段包括音频文件路径、转录文本和说话者ID。数据分割显示训练集有6622个样本,测试集有130个样本。

ClarinPL Sejm/Senat Speech Corpus dataset contains 97 hours of Polish parliamentary speech recordings, which were published on the ClarinPL website. Each data point in the dataset typically includes the audio file path (commonly referred to as `file`) and its transcribed text (referred to as `text`). The audio files are in Polish. The structure of the dataset includes data instances, data fields and data splits. An example data instance shows the audio file path, ID, speaker ID and transcribed text. The data fields include the audio file path, transcribed text and speaker ID. The data splits indicate that the training set contains 6622 samples while the test set has 130 samples.

提供机构:
jimregan
原始信息汇总

数据集概述

数据集名称

  • 名称: ClarinPL Sejm/Senat Speech Corpus

数据集描述

  • 摘要: 该数据集包含97小时的议会演讲,这些演讲发布在ClarinPL网站上。
  • 语言: 音频为波兰语。

数据集结构

  • 数据实例: 每个数据点包括音频文件路径(通常称为file)和其转录文本(称为text)。
  • 数据字段:
    • file: 下载的音频文件(.wav格式)的路径。
    • text: 音频文件的转录文本。
    • speaker_id: 音频发言者的ID。
  • 数据分割:
    • 训练集: 6622个实例
    • 测试集: 130个实例

数据集创建

  • 数据来源: 原始数据
  • 注释创建者: 专家生成

使用数据集的考虑

  • 社会影响: 需要更多信息
  • 偏见讨论: 需要更多信息
  • 其他已知限制: 需要更多信息

附加信息

  • 许可证: 其他
  • 多语言性: 单语种
  • 大小类别: 1K<n<10K
  • 任务类别: 其他, 自动语音识别
二维码
社区交流群
二维码
科研交流群
商业服务