sulabhkatiyar/ne-asr-trp-aug
收藏资源简介:
NE ASR增强数据集——Kokborok(trp)是一个用于自动语音识别(ASR)的增强数据集,专门针对Kokborok语言(ISO 639-3代码:trp)。Kokborok是一种藏缅语系语言,在印度特里普拉邦使用,属于声调语言,具有高/低两种音位声调。该数据集基于原始数据集sulabhkatiyar/ne-asr-trp(来自ARTPARK-IISc Vaani项目)进行增强,当前版本为v2(重建于2026-05-29),仅应用速度扰动增强(因子为0.9、1.0、1.1),避免了音高移位,以保护声调对比。数据集包含训练集(7,533个样本,3倍增强)、验证集(274个样本)和测试集(279个样本),估计原始音频时长约3.8小时,增强后约11.4小时。数据格式包括16kHz单声道WAV音频(以Parquet格式存储)、文本转录、语言和增强标签(如original、speed_0.9、speed_1.1)。许可证为CC-BY-4.0。
The NE ASR Augmented Dataset -- Kokborok (trp) is an augmented automatic speech recognition (ASR) dataset for the Kokborok language (ISO 639-3: trp). Kokborok is a Tibeto-Burman language spoken in Tripura, India, and is a tonal language with two phonemic tones (high/low). The dataset is augmented from the original dataset sulabhkatiyar/ne-asr-trp (from the ARTPARK-IISc Vaani project). The current version is v2 (rebuilt on 2026-05-29), applying only speed perturbation augmentation (factors: 0.9, 1.0, 1.1) while disabling pitch shift to preserve tonal contrasts. It includes a training set (7,533 samples, 3x augmentation), a validation set (274 samples), and a test set (279 samples), with an estimated original duration of ~3.8 hours and augmented duration of ~11.4 hours. The data format consists of 16kHz mono WAV audio (stored as Parquet), text transcriptions, language, and augmentation labels (e.g., original, speed_0.9, speed_1.1). The license is CC-BY-4.0.




