Event-Conditioned Causal Extraction in Saudi-Dialect: A Comparative Study of Dialect-Trained BERTs and LLM Prompting
收藏资源简介:
This dataset contains annotated Arabic tweets related to sick-leave discussions in Saudi social-media contexts. The dataset was developed to support research on causality extraction from informal Arabic text, with a focus on identifying whether a reason or cause for sick leave is expressed and how that cause is linguistically represented. Each tweet was annotated for multiple causality-related elements, including: cause presence, cause span text, cause category, causal marker presence, and causal marker text. Cause categories include health-related reasons, well-being, caregiving, lifestyle or social reasons, and general or unspecified causes. The dataset is intended to support evaluation of Arabic natural language processing models, particularly large language models and Arabic/dialect-specific transformer models, on classification and span extraction tasks. The tweets reflect informal, user-generated Arabic text and may include Saudi dialectal expressions, spelling variation, abbreviations, implicit reasoning, and non-standard grammar.



