遇见数据集

rughimire/slr54nepali-curated

收藏
Hugging Face2026-05-19 更新2026-05-31 收录
官方服务:

资源简介:

Curated Open SLR Nepali Dataset (SLR54 - Nepali) 是一个基于OpenSLR SLR54尼泊尔语语音数据集的整理版本,专为学术实验设计。该数据集语言为尼泊尔语,来源是Open SLR (SLR54)的本地整理子集。数据集包含train.tsv、valid.tsv和test.tsv文件,这些文件以制表符分隔,字段包括file_id、speaker_id和transcript。原始数据集未提供推荐的数据分割,因此这个整理版本提供了标准的训练集、验证集和测试集划分,便于公平评估和比较相关研究。数据集的音频资源托管在Hugging Face平台,而分割文件存储在本仓库中。数据集统计信息显示,训练集有118,256个样本,验证集和测试集各有约14,774和14,771个样本,所有集合均来自527个独特的说话人,音频时长范围从0.60秒到27.61秒,平均约3秒,转录文本长度从5到128个字符,平均约18个字符。总音频时长约为99.17小时(训练集)、12.37小时(验证集)和12.40小时(测试集)。

The Curated Open SLR Nepali Dataset (SLR54 - Nepali) is a curated version of the OpenSLR SLR54 Nepali speech dataset, prepared for academic experiments. The dataset is in the Nepali language and sourced from a locally curated subset of Open SLR (SLR54). It includes train.tsv, valid.tsv, and test.tsv files, which are tab-separated and contain fields such as file_id, speaker_id, and transcript. The original dataset did not provide recommended train/validation/test splits, so this curated version offers standard splits to enable fair evaluation and comparison of work using the dataset. The audio resources are hosted on Hugging Face, while the split files are stored in this repository. Dataset statistics indicate that the training set has 118,256 samples, the validation set has 14,774 samples, and the test set has 14,771 samples, all from 527 unique speakers. Audio durations range from 0.60 seconds to 27.61 seconds, with an average of about 3 seconds, and transcript lengths range from 5 to 128 characters, with an average of about 18 characters. The total audio duration is approximately 99.17 hours for training, 12.37 hours for validation, and 12.40 hours for testing.

提供机构:
rughimire
二维码
社区交流群
二维码
科研交流群
商业服务