igbosyncorp-igbo-asr-benchmark
收藏资源简介:
该数据集是一个多模态语音数据集,专注于语音的声学、文本及语言学标注。数据集中每个样本包含一段采样率为16kHz的音频录音及其对应的文本句子和音标转写。此外,数据集提供了丰富的语言学标注信息,包括音调类别、形态类别、长度类别和方言信息。每个音频片段都带有精确的时间戳和时长,并附有注释说明。数据集还包含量化统计特征,如高音调计数、低音调计数、降阶计数、词数、平均及最大语素数/词以及平均及最大音节数/词。数据集包含一个训练集,共17个样本,总大小约为3.2 MB。该数据集适用于语音处理、计算语言学、音调分析、方言研究以及语音-文本对齐等任务。
This dataset is a multimodal speech dataset focusing on acoustic, textual, and linguistic annotations of speech. Each sample in the dataset includes an audio recording with a sampling rate of 16kHz, along with its corresponding text sentence and phonetic transcription. Additionally, the dataset provides rich linguistic annotation information, including tone category, morphological category, length category, and dialect information. Each audio segment is accompanied by precise timestamps and duration, with annotation notes. The dataset also includes quantitative statistical features such as high tone count, low tone count, downstep count, word count, average and maximum morphemes per word, and average and maximum syllables per word. The dataset contains a training set with 17 samples, totaling approximately 3.2 MB. It is suitable for tasks such as speech processing, computational linguistics, tone analysis, dialect research, and speech-text alignment.
数据集概述
数据集名称:IgboSyncorp-Igbo ASR Benchmark
数据集地址:https://huggingface.co/datasets/str20tbl/igbosyncorp-igbo-asr-benchmark
数据集描述
该数据集是一个用于伊博语(Igbo)自动语音识别(ASR)基准测试的数据集,包含音频文件及其对应的文本标注和丰富的语言学特征。
数据特征
数据集包含22个字段,涵盖音频、文本、语言学标注和统计信息:
- 音频:
audio(采样率:16000 Hz) - 文本标注:
sentence(句子文本)、phonetic(音标)、file_id(文件ID)、annotation_id(标注ID) - 语言学分类:
tone_category(声调类别)、morph_category(形态类别)、length_category(长度类别)、dialect(方言)、notes(备注) - 时间信息:
start_ms(开始毫秒)、end_ms(结束毫秒)、duration_s(持续时间,秒) - 声调统计:
n_high(高声调数量)、n_low(低声调数量)、n_downstep(降阶调数量) - 词汇与语素统计:
n_words(单词数量)、avg_morphemes_per_word(平均每词语素数)、max_morphemes_per_word(最大每词语素数) - 音节统计:
avg_syllables_per_word(平均每词音节数)、max_syllables_per_word(最大每词音节数)
数据集划分
- 训练集(train):共17个样本,数据大小约3.22 MB(3,215,843 字节)
配置信息
- 默认配置名称:
default - 数据文件路径:
data/train-*(训练集)




