遇见数据集

sulabhkatiyar/ne-asr-grt-aug

收藏
Hugging Face2026-05-19 更新2026-05-31 收录
官方服务:

资源简介:

NE ASR增强数据集——Garo(grt)是一个用于自动语音识别的增强数据集,专门针对Garo语言(ISO 639-3代码:grt),这是一种在印度梅加拉亚邦使用的藏缅语系非声调语言。该数据集从原始数据集sulabhkatiyar/ne-asr-grt增强而来,原始数据源自ARTPARK-IISc Vaani项目。原始数据包含约47.3小时的音频(33,480个训练样本),通过应用速度扰动(0.9倍和1.1倍)和音高移位(-2、-1、+1、+2半音)等转换,将训练样本增强至234,360个(7倍增强),总时长约331.3小时。数据集还包括4,253个验证样本和4,101个测试样本。音频数据以16kHz单声道WAV格式存储为Parquet文件,包含音频字节、文本转录、语言和增强标签(如“original”、“speed_0.9”、“pitch_-2”等)。该数据集旨在支持低资源语言的语音识别研究,遵循CC-BY-4.0许可。

The NE ASR Augmented Dataset -- Garo (grt) is an augmented automatic speech recognition dataset for the Garo language (ISO 639-3: grt), a Tibeto-Burman language spoken in Meghalaya, India. It is augmented from the original dataset sulabhkatiyar/ne-asr-grt, which contains transcribed speech data from the ARTPARK-IISc Vaani project. The original data includes approximately 47.3 hours of audio (33,480 training samples), and through transformations such as speed perturbation (0.9x and 1.1x) and pitch shift (-2, -1, +1, +2 semitones), the training samples are augmented to 234,360 samples (7x augmentation), with an estimated total duration of about 331.3 hours. The dataset also includes 4,253 validation samples and 4,101 test samples. Audio data is stored as 16kHz mono WAV in Parquet format, containing audio bytes, text transcriptions, language, and augmentation labels (e.g., original, speed_0.9, pitch_-2). This dataset is designed to support speech recognition research for low-resource languages and is licensed under CC-BY-4.0.

提供机构:
sulabhkatiyar
二维码
社区交流群
二维码
科研交流群
商业服务