quran_nemotron_dual_v5_npl
收藏资源简介:
该数据集名为 Quran Nemotron Dual V5 Assets,是用于阿拉伯语古兰经(Hafs 诵读风格)双输出训练管线的重现性资产。数据集包含两个独立存储的目标:声学 RNNT 目标和基于 Quran-Lab 的 `tj1` 音素 CTC 目标,分别用于不同的训练任务。数据集中每个样本都记录了其来源仓库和固定修订版本,以确保可追溯性。数据集遵循 NPL-1.1(非营利/共享许可)及源特定许可条款。`bootstrap_v4` 目录包含经过验证、说话人无重叠的 v4 语料库(已转换为 Schema 版本 5)。电话录音及其他新增数据源仅在通过访问、许可、去重和泄漏检测后的安全条件下发布。评估基准音频从不包含在训练分片中,以保证评估的公正性。
This dataset is named Quran Nemotron Dual V5 Assets, which is a reproducible asset for dual output training pipeline of the Arabic Quran (Hafs recitation style). It contains two independently stored targets: acoustic RNNT targets and tj1 phoneme CTC targets based on Quran-Lab, for different training tasks. Each sample records its source repository and fixed revision for traceability. The dataset is licensed under NPL-1.1 (Non-Profit/Share Alike) and source-specific licenses. The bootstrap_v4 directory contains a validated, speaker-disjoint v4 corpus (converted to Schema version 5). Telephone recordings and other additional data sources are only released under security conditions after access, licensing, deduplication, and leakage detection. Evaluation benchmark audio is never included in training splits to ensure unbiased evaluation.
数据集概述
该数据集为《古兰经》尼莫创双输出模型的复现资产,专注于阿拉伯语哈夫斯(Hafs)古兰经的语音训练流水线。
核心内容
- 双输出目标:包含基于声学 RNNT 的预测目标和基于古兰经实验室
tj1音素 CTC 的预测目标,两者独立存储。 - 数据来源:每条数据记录均包含源仓库及固定的修订版本号,不授予源录音的新许可,需遵循源特定条款及古兰经实验室 NPL-1.1 非营利/共享署名限制。
- 数据版本:
bootstrap_v4目录为经过验证的说话人分离 v4 语料库,已转换为 schema 版本 5。新增的电话录音等数据仅在访问、许可、去重和泄漏检查通过后发布。 - 评估安全:评估基准音频从不包含在训练分片(shards)中。




