Noor-Sharaye v.1. A Benchmark Dataset of Complex Words for Evaluating Arabic Analyzers
收藏官方服务:
资源简介:
The Noor-Sharaye dataset is a morphologically annotated Classical Arabic corpus containing approximately 205,000 word instances extracted from 17 different Classical Arabic books, including Quranic, Fiqh, Hadith, and historical texts. Each token is enriched with detailed linguistic annotations such as Stem Lemma Root part-of-speech tags (pos) Segmentation Grammatical case Gender, Number Affix-level features The data are encoded in UTF-8 XLS, XML, and JSON formats for broad compatibility. This resource supports stemming, root extraction, morphological analysis, and benchmarking of AI-based models in Arabic Natural Language Processing
提供机构:
Zenodo创建时间:
2026-07-21



