CARDAMOM
收藏资源简介:
CARDAMOM是一个面向阿拉伯语微方言的语音数据集,由多所大学研究者联合创建,包含约40小时转录语音,覆盖埃及、约旦、黎巴嫩、毛里塔尼亚、巴勒斯坦、沙特六国21种微方言,共计27,420条话语。数据来自YouTube非脚本化自然对话,由母语者标注,提供微方言标签、语码转换信息及感知性别。数据集旨在支持自动语音识别(ASR)系统的细粒度评估与适配,解决传统粗粒度方言标签无法捕捉的本地化变异问题。
CARDAMOM is a speech dataset targeting Arabic micro-dialects, jointly created by researchers from multiple universities. It contains approximately 40 hours of transcribed speech, covering 21 micro-dialects across six countries including Egypt, Jordan, Lebanon, Mauritania, Palestine, and Saudi Arabia, with a total of 27,420 utterances. The data is sourced from unscripted natural conversations on YouTube, annotated by native speakers, and provides micro-dialect labels, code-switching information, and perceived gender. This dataset is designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems, addressing the localized linguistic variation issues that cannot be captured by traditional coarse-grained dialect labeling.

- 1CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR不列颠哥伦比亚大学; 哈马德·本·哈利法大学; 比尔宰特大学; 贝鲁特美国大学; 阿拉伯政策研究中心; 法赫德国王石油矿产大学; 塔伊巴大学; 高等数字学院; 伊玛目阿卜杜勒拉赫曼·本·费萨尔大学; 北部边境大学; 约旦科技大学; 巴德尔开罗大学; 埃尔塞维迪科技大学; 帝国理工学院; 黎巴嫩大学; 阿尔贝特大学 · 2026年




