L-HSAB
收藏资源简介:
L-HSAB是首个阿拉伯语黎凡特地区的仇恨言论和辱骂语言数据集,包含5,846条来自叙利亚和黎巴嫩的政治推文,标记为正常、辱骂或仇恨。数据集通过Twitter API收集,重点关注政治敏感话题,如难民、女性、阿拉伯人等,并由三名黎凡特语标注者进行标注。
L-HSAB is the first dataset of hate speech and abusive language in the Levantine Arabic dialect, comprising 5,846 political tweets from Syria and Lebanon, labeled as normal, abusive, or hateful. The dataset was collected via the Twitter API, focusing on politically sensitive topics such as refugees, women, and Arabs, and was annotated by three native Levantine Arabic speakers.
L-HSAB Dataset Summary
Dataset Overview
- Name: L-HSAB (Levantine Hate Speech and Abusive) Dataset
- Description: The first Arabic Levantine Hate Speech and Abusive Language Dataset, proposed in the 3rd Workshop ALW-2019 co-located with ACL-2019.
- Content: 5,846 Syrian/Lebanese political tweets labeled as normal, abusive, or hate.
- Timeframe: Tweets collected between March 2018 and February 2019.
Data Collection
- Method: Tweets scraped via Twitter API (Tweepy) using keywords related to potential targets of abusive/hate speech.
- Sources: User timelines of verified or high-follower count politicians, activists, and TV anchors.
Data Annotation
- Annotators: 3 Levantine-speaking annotators.
- Categories:
- Normal: No offensive content.
- Abusive: Contains offensive, aggressive, insulting, or profanity content.
- Hate: Contains abusive language directed at a specific person or group, demeaning or dehumanizing based on identity.
- Guidelines: Provided with nicknames used in hate/abusive contexts for political parties and groups.
Annotation Evaluation
- Measures:
- Pairwise Percent Agreement Measure (PRAM): 87.24%
- Cohens Kappa (K): 75.8%
- Krippendorff’s Alpha (α): 76.5%
Classification Experiments
- Binary Classification (Normal, Abusive):
- Best model: Naive Bayes
- F-measure: 89.6%
- Multi-Class Classification (Normal, Abusive, Hate):
- Best model: Naive Bayes
- F-measure: 74.4%
Paper Citation
@inproceedings{mulki2019hsab, title={L-HSAB: A Levantine Twitter Dataset for Hate Speech and Abusive Language}, author={Mulki, Hala and Haddad, Hatem and Ali, Chedi Bechikh and Alshabani, Halima}, booktitle={Proceedings of the Third Workshop on Abusive Language Online}, pages={111--118}, year={2019} }




