遇见数据集

NOOR dataset

收藏
Zenodo2026-04-19 更新2026-05-26 收录
官方服务:

资源简介:

A Classical Arabic dataset for stemming and morphological analysis containing approximately 260,000 word–stem pairs. Only a subset of 7,042 word–stem pairs (derived from the Quran) is publicly available. The data is formatted for supervised learning and provided in structured formats such as CSV and XML,PDF supporting character-level sequence-to-sequence NLP models. The dataset can be used for Arabic stemming, lemmatization, morphological analysis, root extraction, and evaluation of deep learning models in Classical Arabic NLP.

提供机构:
Zenodo
创建时间:
2026-04-19
二维码
社区交流群
二维码
科研交流群
商业服务