ARAPROPWIKID3K
收藏资源简介:
ARAPROPWIKID3K数据集是一个包含3362对阿拉伯维基百科专有名词及其英文对应解释的数据集,其中专有名词已经经过人工标注,带有完整的词尾变化和词干化信息。数据集旨在解决阿拉伯维基百科中常见的问题,即专有名词的发音和解释存在歧义,因为阿拉伯语的正字法通常省略了变音符号。该数据集为阿拉伯语自然语言处理(NLP)中的三个关键任务——转写、变音和词干化——提供了一个基准,有助于研究这些任务的交集。数据集的内容涵盖了各种来源的专有名词,包括人名、地名和组织名,并反映了基于解释的多个有效变音。该数据集的创建过程包括从ARAPROP数据集中随机选择3,000个独特的阿拉伯语专有名词进行人工标注,并形成了3,362个阿拉伯语-英文解释对。数据集的应用领域包括提高阿拉伯语NLP模型的性能,特别是在专有名词的转写和变音方面,以及促进对阿拉伯语专有名词资源的研究。
The ARAPROPWIKID3K dataset is a collection of 3,362 pairs of Arabic Wikipedia proper nouns and their corresponding English glosses, where the proper nouns have been manually annotated with complete inflection and stemming information. This dataset aims to address a common issue in Arabic Wikipedia: ambiguity in the pronunciation and interpretation of proper nouns, as Arabic orthography typically omits diacritics. It serves as a benchmark for three core tasks in Arabic natural language processing (NLP): transliteration, diacritization, and stemming, facilitating research into the intersection of these tasks. The dataset covers proper nouns from various sources, including person names, place names, and organizational names, and reflects multiple valid diacritic variants based on their English explanations. The construction process of this dataset involves randomly selecting 3,000 unique Arabic proper nouns from the ARAPROP dataset for manual annotation, resulting in 3,362 Arabic-English gloss pairs. The application scenarios of this dataset include improving the performance of Arabic NLP models, particularly in the transliteration and diacritization of proper nouns, as well as facilitating research on Arabic proper noun resources.

- 1Proper Name Diacritization for Arabic Wikipedia: A Benchmark Dataset纽约大学阿布扎比分校、埃及艾因沙姆斯大学、纽约州立大学石溪分校、马耳他大学人工智能系 · 2025年



