Arabic POS & Lemma Evaluation Corpus: Stanza, Farasa, and CAMeL Tools on Classical and Modern Texts
收藏资源简介:
A human-evaluated corpus for benchmarking three Arabic NLP toolkits (Stanza, Farasa, CAMeL Tools) on POS tagging and lemmatisation. Covers 10 texts (5 classical, 5 modern) across 10 subject domains. Contains 60 per-system, per-task, per-text TSV files, 2 long-format master TSVs (≈79K rows), source texts, pipeline notebooks, an 18-category canonical tagset specification with bottom-up mapping rules, and a tested canonicalisation script.
本数据集为一款人工评估语料库,用于在词性标注(Part-of-Speech Tagging,POS tagging)与词形还原(lemmatisation)任务上,对三款阿拉伯语自然语言处理工具包(Stanza、Farasa、CAMeL Tools)开展基准测试。该语料库涵盖10篇文本(含5篇古典阿拉伯语文本、5篇现代阿拉伯语文本),覆盖10个学科领域。数据集包含60个针对各工具包、各任务、各文本的制表符分隔值文件(Tab-Separated Values,TSV)、2个长格式主TSV文件(约79000行)、原始源文本、流水线处理笔记、1份包含自底向上映射规则的18类规范标注集(canonical tagset)说明文档,以及1份经过验证的规范化处理脚本。



