遇见数据集

IPA Transcriptions of the Arabic Speech Corpus (ASC)

收藏
Zenodo2025-09-13 更新2026-05-26 收录
官方服务:

资源简介:

This dataset provides International Phonetic Alphabet (IPA) transcriptions for the Arabic Speech Corpus (ASC; Halabi, 2016), a ∼4-hour collection of South Levantine Arabic (Damascian accent) speech originally designed for speech synthesis. The ASC consists of high-quality studio recordings read by a single native speaker, with transcripts derived from Al Jazeera Learn materials and automatically generated sentences optimized for phonetic coverage. The original ASC release contained orthographic transcriptions and phonetic annotations in Buckwalter format. For this dataset, we constructed a Buckwalter-to-IPA conversion module to systematically map the provided Buckwalter transcriptions into IPA. The conversion relied on the rules and lookup tables described in Halabi (2016), supplemented with Arabic romanization standards. Citation If you use this dataset in your work, please cite: Jacob Lee Suchardt, Hana El-Shazli, Pierluigi Cassotti. 2025. Towards Language-Agnostic STIPA: Universal Phonetic Transcription to Support Language Documentation at Scale. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing.

提供机构:
Zenodo
创建时间:
2025-09-13
二维码
社区交流群
二维码
科研交流群
商业服务