CV-18 NER
收藏资源简介:
CV-18 NER是由ELYADATA团队创建的首个阿拉伯语语音命名实体识别数据集,基于Common Voice 18阿拉伯语子集构建。该数据集包含约8小时15分钟的语音数据,共计7119条标注样本,采用细粒度的Wojood标注体系(21种实体类型)。数据集通过自动预标注与人工修订相结合的方式构建,首先利用AraBERT模型生成伪标签,再由专业标注员进行人工校正,最终过滤掉不含实体的语句。该数据集主要用于评估端到端阿拉伯语语音命名实体识别系统的性能,解决阿拉伯语语音理解任务中标注资源匮乏的问题,为低资源场景下的语义解析提供基准。
CV-18 NER is the first Arabic speech named entity recognition (NER) dataset created by the ELYADATA team, constructed based on the Arabic subset of Common Voice 18. This dataset contains approximately 8 hours and 15 minutes of speech data, with a total of 7119 annotated samples, and adopts the fine-grained Wojood annotation schema with 21 entity types. It is built through a combined workflow of automatic pre-annotation and manual revision: first, the AraBERT model is used to generate pseudo-labels, then professional annotators perform manual correction, and finally sentences without any entities are filtered out. This dataset is mainly used to evaluate the performance of end-to-end Arabic speech named entity recognition systems, addressing the problem of scarce annotated resources in Arabic speech understanding tasks, and providing a benchmark for semantic parsing in low-resource scenarios.
CV-18 NER 数据集概述
数据集基本信息
- 数据集名称:CV-18 NER
- 许可证:cc-by-nc-4.0
- 任务类别:自动语音识别
- 语言:阿拉伯语(ar)
- 标签:命名实体识别、语音命名实体识别
数据集描述
CV-18 NER 是首个公开可用的、用于从阿拉伯语语音中进行命名实体识别的数据集。该数据集通过为阿拉伯语 Common Voice 18 语料库添加手动命名实体识别标注而创建,标注遵循细粒度的 Wojood 模式,涵盖 21 种实体类型。
该数据集为评估流水线系统(自动语音识别 + 文本命名实体识别)和端到端语音命名实体识别模型提供了基准。对于低资源环境和形态复杂语言(如阿拉伯语)的研究具有重要价值。
更多信息
详细信息可查阅论文:CV-18 NER: Augmented Common Voice for Named Entity Recognition from Arabic Speech。

- 1CV-18 NER: Augmented Common Voice for Named Entity Recognition from Arabic SpeechELYADATA · 2026年



