quran-audio
收藏资源简介:
QuranLab — Quran Recitation Audio 是一个专注于古兰经诵读音频的参考层数据集。它不直接包含音频文件,而是提供了一个诚实许可的元数据框架,包括音频的外部链接和词语级时间戳。数据集的核心组成部分包括:1) 一个规范的诵经人分类法,涵盖245位来自不同时代、地区和风格的诵经人,并标注了其来源、诵读方式(如Hafs、Warsh)和风格(如murattal、mujawwad);2) 一个按节(ayah)的参考清单,包含47个录音,覆盖古兰经全部6,236节,每行数据提供对应的音频URL;3) 一个按章(surah)的参考层,包含来自mp3quran的285个录音,覆盖15种不同的诵读方式,提供整章音频的URL。其中,33个按节录音提供了精确到词语级别的时间戳(`segments`字段),时间戳以CC-BY-4.0许可发布。这些时间戳部分来自已有的开源项目(cpfair/quran-align),部分由数据集创建者使用Apache-2.0许可的声学模型和torchaudio通过强制对齐技术生成并验证。数据集设计用于自动语音识别(ASR)、音频对齐、韵律分析等任务,并可作为文本语料库`quranlab/quran`的配套音频参考。用户需注意,数据集仅提供元数据和链接,音频文件需从原始CDN获取并遵守其条款;对于非Hafs(如Warsh)的录音,需注意其节编号可能与标准Hafs 6,236节体系存在差异。
QuranLab — Quran Recitation Audio is a reference layer dataset focused on Quran recitation audio. It does not directly contain audio files but provides a metadata framework with honest licensing, including external links to audio and word-level timestamps. The core components of the dataset include: 1) A canonical taxonomy of reciters, covering 245 reciters from different eras, regions, and styles, annotated with their sources, recitation methods (riwayah, such as Hafs, Warsh), and styles (e.g., murattal, mujawwad); 2) A verse (ayah) reference manifest containing 47 recordings covering all 6,236 verses of the Quran, with each row providing the corresponding audio URL; 3) A chapter (surah) reference layer containing 285 recordings from mp3quran, covering 15 different recitation methods, providing URLs for full chapter audio. Among these, 33 verse recordings provide precise word-level timestamps (the `segments` field), released under the CC-BY-4.0 license. These timestamps are partly from existing open-source projects (cpfair/quran-align) and partly generated and verified by the dataset creators using Apache-2.0 licensed acoustic models and torchaudio through forced alignment techniques. The dataset is designed for tasks such as automatic speech recognition (ASR), audio alignment, and prosody analysis, and serves as a complementary audio reference for the text corpus `quranlab/quran`. Users should note that the dataset only provides metadata and links; audio files must be obtained from the original CDN and comply with its terms. For non-Hafs recordings (e.g., Warsh), note that the verse numbering may differ from the standard Hafs 6,236-verse system.
数据集概述
QuranLab — Quran Recitation Audio 是一个专为《古兰经》诵读音频设计的参考数据集,不直接提供音频文件,而是提供统一的引诵者分类法、经文节级别的引用清单以及开放许可的词级时间戳。
核心特性
- 语言: 阿拉伯语(
ar) - 引诵者: 共 245 位,提供阿拉伯语和英语名称、国籍、时代、流行度等级以及可用的诵读风格(riwāyāt)
- 数据集大小: 记录数在 100,000 到 1,000,000 之间
- 许可协议:
- 词级时间戳:CC-BY-4.0
- 音频:仅提供引用(链接到原始来源,音频版权归引诵者和制作人所有)
- 配套文本语料库:
quranlab/quran(可通过verse_key字段连接)
数据集结构
数据集包含多个配置(config),可通过 load_dataset("quranlab/quran-audio", "config_name") 加载。
| 配置名称 | 说明 |
|---|---|
manifest(默认) |
节级引用索引:包含 47 份录音(来自 everyayah.com),每节经文对应一个外部 audio_url,覆盖标准的 6,236 节经文 |
reciters |
统一的 245 位引诵者分类法,每行一个引诵者 |
surah-manifest |
章级广度层:285 份录音(来自 mp3quran.net),涵盖 15 种 qirāʾāt / riwāyāt |
<recitation_id> |
每份节级录音的单独配置(共 47 个,如 husary、mishary-alafasy),每份 6,236 行;其中 44 份包含 segments 词级时间戳 |
主要字段(节级录音配置)
verse_key:标准surah:ayah键,用于与quranlab/quran连接surah、ayah:章号和节号recitation_id、reciter_id:录音和引诵者标识符name_en、name_ar:引诵者英文和阿拉伯语名称riwayah、style:诵读方式(如hafs-asim)和风格(murattal/mujawwad/muallim)audio_url:指向原始音频的外部链接has_word_timing:是否存在开放许可的词级时间戳segments:每词的起始和结束时间(毫秒),包括word_position、start_ms、end_mstiming_source、timing_license、timing_attribution:时间戳的来源和许可证信息source、source_url、license、attribution:被引用音频的来源和条款
数据贡献
- 统一的引诵者分类法: 识别并整理 245 位引诵者,解决跨来源的名称不一致问题
- 新增 33 份开放许可的词级时间戳: 在上游项目
cpfair/quran-align已提供的 11 份基础上,额外生成 33 份,使总数达到 44 份 - 清晰的逐行许可标注: 每一行都带有独立的许可证和归属信息,音频仅引用不重新托管
数据集来源
- 节级音频引用: everyayah.com
- 章级音频引用: mp3quran.net
- 上游时间戳(11 份录音):
cpfair/quran-align(CC-BY-4.0) - 对齐参考文本: Tanzil 项目(Uthmani 文本)
直接用途
- 自动语音识别(ASR):提供节对齐的参考和词级时间戳,用于训练和评估
- 文本转语音(TTS) 与韵律/发音研究




