Q-CSMP v2: A Minimal-Pair Pilot Dataset for Contextual Sense Disambiguation in Qur'anic Arabic
收藏资源简介:
Q-CSMP v2 (Qur'anic Contextual Sense Minimal Pairs, version 2) — a pilot dataset for contextual sense disambiguation in Qur'anic Arabic. Built 2026-09-28. Grade: Verified (pilot) — independent replication run byte-identical (surface_results.json md5 3367e8630c6167b2c18c65d042ca5cd2). Scale: 5,769 base items (23 audited + 5,746 keyword-rule) × 5 variants = 28,841 records (4 root-substitutions skipped). 48 lemmas (9 v1 + 39 auto-mined), 115 lemma-senses (84 attested in test). Splits (whole-surah holdout): train 4,691 / dev 451 / test 627. Address match: MASAQ 5,768/5,769 · Tanzil 5,769/5,769. Honest pilot framing — NOT a finished benchmark: Labels are keyword-rule over English glosses (label_source=keyword-rule); 38/95 sense groups are heuristic-flagged suspect. Not ground truth — expert adjudication (inter-annotator agreement κ ≥ 0.61) is pending and is a hard prerequisite before any peer-reviewed publication of the associated paper. Scale target of ≥10k base items was MISSED — honest ceiling ≈ 5.8k. Says nothing about trained-LLM behavior (linear models, 48 lemmas, zero pre-training). Do not use for theological claims. Measured results (test n=627): M_surface accuracy 0.7656 / mean sense recall 0.7421; M_Q accuracy 0.7129 / recall 0.6866. Δ_Q pilot: −0.0555 (mean sense recall) — surface statistics beat root/pattern/morph features at this scale with linear models (negative result, recorded as such). Files: qcsmp_v2.csv, qcsmp_v2.jsonl (28,841 records), and DATASET_CARD.md (full dataset card with schema, provenance, and build details).



