遇见数据集

Noor-Sharaye v.1. A Benchmark Dataset of Complex Words for Evaluating Arabic Analyzers

收藏
Zenodo2026-07-21 更新2026-08-01 收录
官方服务:

资源简介:

The Noor-Sharaye dataset is a morphologically annotated Classical Arabic corpus containing approximately 205,000 word instances extracted from 17 different Classical Arabic books, including Quranic, Fiqh, Hadith, and historical texts. Each token is enriched with detailed linguistic annotations such as Stem Lemma Root part-of-speech tags (pos) Segmentation Grammatical case Gender, Number Affix-level features The data are encoded in UTF-8 XLS, XML, and JSON formats for broad compatibility. This resource supports stemming, root extraction, morphological analysis, and benchmarking of AI-based models in Arabic Natural Language Processing

提供机构:
Zenodo
创建时间:
2026-07-21
二维码
社区交流群
二维码
科研交流群
商业服务