Arabic Segmentation, Lemmatization, and POS Evaluation Dataset: Stanza, Farasa, and CAMeL Tools on Classical and Modern Texts
收藏资源简介:
Case-level evaluation data for three Arabic NLP systems — Stanza, Farasa, and CAMeL Tools — on a purposively assembled 10,000-word Arabic corpus of ten texts: five classical and five modern, across ten subject areas. The package contains frozen case tables for three task families: segmentation (44,218 cases), part-of-speech tagging (44,221 cases), and lemmatization/stemming (33,474 cases); Farasa's output is reported as stemming, not lemmatization. Each case carries the system prediction, the manually corrected reference, the verdict (T, F, T/F for cases affected by differing unit boundaries), the canonical mapping used for cross-system comparison, and full provenance back to the source workbook, sheet, row, and checksum. Deterministic summary tables reproduce every figure reported in the dissertation, and scripts/rebuild_results.py and scripts/verify_package.py regenerate and verify them from the frozen cases. The corpus is a diagnostic comparison corpus, not a probability sample of Arabic; results are descriptive, with no significance tests or confidence intervals. Systems were run with their released settings; none was retrained or tuned, and the corpus was used for evaluation only. بيانات التقييم على مستوى الحالة لثلاثة أنظمة لمعالجة العربية آليًّا: ستانزا وفراسة وأدوات كامل، على مدونة قصدية من عشرة نصوص بواقع عشرة آلاف كلمة، خمسة نصوص قديمة وخمسة حديثة في عشرة موضوعات. وتشمل الحزمة حالات التقطيع والتوسيم النحوي والتأصيل المعجمي والتجذيع، مع التصحيح اليدوي المرجعي، وسلسلة التتبع إلى ملف المصدر وسطره وبصمته، وملخصات محسوبة آليًّا تطابق ما ورد في الرسالة، وسكربتَي إعادة البناء والتحقق



