遇见数据集

Arabic Segmentation, Lemmatization, and POS Evaluation Dataset: Stanza, Farasa, and CAMeL Tools on Classical and Modern Texts

收藏
Zenodo2026-09-30 更新2026-10-01 收录
官方服务:

资源简介:

Case-level evaluation data for three Arabic NLP systems — Stanza, Farasa, and CAMeL Tools — on a purposively assembled 10,000-word Arabic corpus of ten texts: five classical and five modern, across ten subject areas. The package contains frozen case tables for three task families: segmentation (44,218 cases), part-of-speech tagging (44,221 cases), and lemmatization/stemming (33,474 cases); Farasa's output is reported as stemming, not lemmatization. Each case carries the system prediction, the manually corrected reference, the verdict (T, F, T/F for cases affected by differing unit boundaries), the canonical mapping used for cross-system comparison, and full provenance back to the source workbook, sheet, row, and checksum. Deterministic summary tables reproduce every figure reported in the dissertation, and scripts/rebuild_results.py and scripts/verify_package.py regenerate and verify them from the frozen cases. The corpus is a diagnostic comparison corpus, not a probability sample of Arabic; results are descriptive, with no significance tests or confidence intervals. Systems were run with their released settings; none was retrained or tuned, and the corpus was used for evaluation only. بيانات التقييم على مستوى الحالة لثلاثة أنظمة لمعالجة العربية آليًّا: ستانزا وفراسة وأدوات كامل، على مدونة قصدية من عشرة نصوص بواقع عشرة آلاف كلمة، خمسة نصوص قديمة وخمسة حديثة في عشرة موضوعات. وتشمل الحزمة حالات التقطيع والتوسيم النحوي والتأصيل المعجمي والتجذيع، مع التصحيح اليدوي المرجعي، وسلسلة التتبع إلى ملف المصدر وسطره وبصمته، وملخصات محسوبة آليًّا تطابق ما ورد في الرسالة، وسكربتَي إعادة البناء والتحقق

提供机构:
Zenodo
创建时间:
2026-09-30
二维码
社区交流群
二维码
科研交流群
商业服务