QuranMorph
收藏资源简介:
QuranMorph数据集是一个形态学标注的《古兰经》语料库,包含77,429个词。该语料库由伯宰特大学的研究团队创建,每个词都经过三位语言学家手动进行词干化和词性标注。词干化过程使用了Qabas阿拉伯语词典数据库中的词干,该数据库与110个词典和2百万个词的语料库相关联。词性标注使用了细粒度的SAMA/Qabas词性标注集,包含了40个标签。QuranMorph语料库是开源的,作为SinaLab资源的一部分公开提供。该数据集旨在解决计算语言学中古典阿拉伯语资源不足的问题,并为阿拉伯语的自然语言处理研究提供支持。
The QuranMorph dataset is a morphologically annotated corpus of the Quran, containing 77,429 words. This corpus was developed by a research team at Birzeit University, where every word has been manually lemmatized and part-of-speech (POS) tagged by three linguists. Lemmatization utilizes lemmas sourced from the Qabas Arabic Lexicon Database, which is associated with 110 dictionaries and a 2-million-word corpus. The POS tagging follows the fine-grained SAMA/Qabas POS tagset, which includes 40 distinct tags. The QuranMorph corpus is open-source and publicly released as part of the SinaLab resources. This dataset is designed to address the scarcity of Classical Arabic resources in computational linguistics and support natural language processing (NLP) research related to Arabic.




