遇见数据集

Arabic POS & Lemma Evaluation Corpus: Stanza, Farasa, and CAMeL Tools on Classical and Modern Texts

收藏
Zenodo2026-05-15 更新2026-06-05 收录
官方服务:

资源简介:

A human-evaluated corpus for benchmarking three Arabic NLP toolkits (Stanza, Farasa, CAMeL Tools) on POS tagging and lemmatisation. Covers 10 texts (5 classical, 5 modern) across 10 subject domains. Contains 60 per-system, per-task, per-text TSV files, 2 long-format master TSVs (≈79K rows), source texts, pipeline notebooks, an 18-category canonical tagset specification with bottom-up mapping rules, and a tested canonicalisation script.

提供机构:
Zenodo
创建时间:
2026-05-11
二维码
社区交流群
二维码
科研交流群
商业服务