speed-tb/testgloss2
收藏资源简介:
Glossing Test 4是一个用于自动语音识别(ASR)任务的数据集,主要包含印地语(hi)的音频文件和对应的文本转录。数据集按训练、测试、验证等分割组织,每个样本包含音频路径、唯一标识符、原始文件名、文本转录、说话者ID、边界ID、开始和结束时间、TextGrid JSON数据等元数据。数据集采用CC-By-NC-SA-4.0许可,允许非商业用途下的使用,商业用途需联系数据集提供者。
Glossing Test 4 is a dataset for automatic speech recognition (ASR) tasks, primarily containing audio files and corresponding text transcriptions in Hindi (hi). The dataset is organized by splits (e.g., train, test, validation), with each sample including metadata such as audio path, unique identifier, original filename, text transcription, speaker ID, boundary ID, start and end times, and TextGrid JSON data. The dataset is licensed under CC-By-NC-SA-4.0, allowing use for non-commercial purposes, while commercial use requires contacting the dataset provider.




