遇见数据集

WhissleAI/multilingual-libri-test-spanish

收藏
Hugging Face2024-07-25 更新2025-04-08 收录
官方服务:

资源简介:

--- dataset_info: features: - name: file dtype: string - name: audio dtype: audio: sampling_rate: 16000 - name: text dtype: string - name: speaker_id dtype: int64 - name: chapter_id dtype: int64 - name: id dtype: string - name: tasks sequence: string - name: instruction dtype: string splits: - name: train num_bytes: 600403524.88 num_examples: 2385 download_size: 603662076 dataset_size: 600403524.88 configs: - config_name: default data_files: - split: train path: data/train-* --- Automated annotation of the LibriSpeech Spanish test set. Entity annotations from OnTo Notes tagger and emotion from an audio emotion classifier.

数据集信息: 特征: - 名称:file,数据类型:字符串 - 名称:audio,数据类型: 音频:采样率为16000Hz - 名称:text,数据类型:字符串 - 名称:speaker_id,数据类型:int64 - 名称:chapter_id,数据类型:int64 - 名称:id,数据类型:字符串 - 名称:tasks,数据类型:字符串序列 - 名称:instruction,数据类型:字符串 数据集划分: - 名称:训练集(train),字节数:600403524.88,样本数量:2385 下载大小:603662076字节 数据集存储大小:600403524.88字节 配置项: - 配置名称:默认配置(default),数据文件: - 划分:train,路径:data/train-* 本数据集为LibriSpeech西班牙语测试集的自动化标注数据,其中实体标注基于OnTo Notes标注器生成,情感标注则通过音频情感分类器得到。

提供机构:
WhissleAI
二维码
社区交流群
二维码
科研交流群
商业服务