MasriSpeech-Full
收藏资源简介:
MasriSpeech-Full是一个大规模的埃及阿拉伯语音数据集,包含52,914个专业标注的音频样本,总时长超过3,100小时。数据集旨在推动方言阿拉伯语的自动语音识别和语音处理研究,支持高质量16kHz语音录制和自然对话风格,适用于自动语音识别、方言研究、语音合成等领域。
MasriSpeech-Full is a large-scale Egyptian Arabic speech dataset comprising 52,914 professionally annotated audio samples with a total duration exceeding 3,100 hours. This dataset aims to advance research on automatic speech recognition (ASR) and speech processing for dialectal Arabic, supports high-quality 16kHz speech recordings and natural conversational styles, and is applicable to domains such as automatic speech recognition, dialect research, and speech synthesis.
MasriSpeech-Full 数据集概述
基本描述
- 数据集名称: MasriSpeech-Full: Large-Scale Egyptian Arabic Speech Corpus
- 数据类型: 语音(音频)与文本转录
- 语言: 埃及阿拉伯语 (arz)、阿拉伯语 (ar)
- 许可证: Apache 2.0
- 发布年份: 2025
- 发布者: Yahya Muhammad Alnwsany
数据集规模
- 总样本数: 52,914
- 训练集: 50,715 样本
- 验证集: 2,199 样本
- 总时长: ~3,100 小时
- 采样率: 16 kHz
- 格式: Parquet
- 数据集大小: 11.57 GB
- 下载大小: 10.26 GB
数据结构
特征字段
- audio: 音频特征对象,包含:
Array: 原始语音波形(1D 浮点数组)Path: 相对音频路径Sampling_rate: 16,000 Hz
- transcription: 埃及阿拉伯语转录文本(字符串)
数据划分
| 划分 | 样本数 | 大小 (GB) | 平均词数 | 空转录 | 非阿拉伯语 |
|---|---|---|---|---|---|
| 训练集 | 50,715 | 10.42 | 13.34 | 6 | 13 |
| 验证集 | 2,199 | 0.36 | 9.60 | 0 | 1 |
语言统计
训练集
- 高频词: في (20,250), و (16,977)
- 高频二元组: (إن, أنا) (1,305)
- 词汇量: 38,451
- 独特说话人: 1,142
验证集
- 高频词: في (519), أنا (412)
- 高频二元组: (شاء, الله) (63)
- 词汇量: 7,892
- 独特说话人: 98
使用方式
加载数据集
python from datasets import load_dataset ds = load_dataset(NightPrince/MasriSpeech-Full, split=train, streaming=True)
预处理示例
python from transformers import Wav2Vec2Processor processor = Wav2Vec2Processor.from_pretrained("facebook/wav2vec2-base-960h")
模型微调
python from transformers import AutoModelForCTC, TrainingArguments model = AutoModelForCTC.from_pretrained("facebook/wav2vec2-base-960h")
引用格式
bibtex @dataset{masrispeech_full, author = {Yahya Muhammad Alnwsany}, title = {MasriSpeech-Full: Large-Scale Egyptian Arabic Speech Corpus}, year = {2025}, publisher = {Hugging Face}, url = {https://huggingface.co/collections/NightPrince/masrispeech-dataset-68594e59e46fd12c723f1544} }
应用场景
- 埃及阿拉伯语自动语音识别 (ASR)
- 方言阿拉伯语语言学研究
- 语音合成与语音克隆
- 低资源语言机器学习模型训练与基准测试




