MultiDiac
收藏资源简介:
MultiDiac是一个多语言数据集,用于评估大型语言模型在阿拉伯语和约鲁巴语中的文本标音效果。该数据集包含多样化的样本,涵盖了各种标音歧义。研究机构 Mohamed Bin Zayed University of Artificial Intelligence 与三位母语为约鲁巴语的语言学家合作收集了约350个基本单词,每个单词都有多个有效的标音形式,并构建了3到4个语境丰富的句子来确保多样性和自然语言歧义。阿拉伯语数据集由一位母语为阿拉伯语的人士和一位L2级熟练人士收集,选择了约42个基本单词,每个单词都有多个标音形式,并构建了3到4个语境不同的句子。
MultiDiac is a multilingual dataset designed to evaluate the text diacritization performance of large language models (LLMs) in Arabic and Yoruba. The dataset includes diverse samples covering various diacritization ambiguities. The Mohamed Bin Zayed University of Artificial Intelligence collaborated with three native Yoruba linguists to collect approximately 350 base words, each with multiple valid diacritization forms, and constructed 3 to 4 context-rich sentences for each word to ensure diversity and natural language ambiguities. The Arabic subset of the dataset was collected by a native Arabic speaker and a proficient second-language (L2) speaker of Arabic, selecting roughly 42 base words each with multiple diacritization forms, and building 3 to 4 sentences with distinct contexts for each word.




