Dataset-of-Urdu-Abusive-Language
收藏资源简介:
公开的乌尔都语辱骂语言数据集,数据集通过使用特定乌尔都语辱骂关键词的Tweepy API抓取推文,并由乌尔都语母语者进行标注。数据集经过预处理和清洗,移除了不需要的元素如停用词、标签、提及和URL。数据集平衡,推文长度在10到256个字符之间。
An open dataset of Urdu abusive language, collected by scraping tweets using the Tweepy API with specific Urdu abusive keywords, and annotated by native Urdu speakers. The dataset has been preprocessed and cleaned, removing unnecessary elements such as stop words, hashtags, mentions, and URLs. The dataset is balanced, with tweet lengths ranging between 10 to 256 characters.
数据集概述
数据集名称
Dataset-of-Urdu-Abusive-Language
数据来源
数据集通过Tweepy API使用特定的乌尔都语辱骂关键词从Twitter上抓取,并由乌尔都语母语者进行标注。
数据预处理
数据集已经过预处理和清洗,移除了停用词、话题标签、提及和URL等不必要元素。数据集是平衡的,推文长度介于10至256个字符之间。
数据集统计
- 总样本数:12071
- 辱骂性推文:5930
- 中性推文:6141
引用信息
若在研究或项目中使用此数据集,请引用以下论文: Khan, A., Ahmed, A., Jan, S., Bilal, M. and Zuhairi, M.F., 2024. Abusive Language Detection in Urdu Text: Leveraging Deep Learning and Attention Mechanism. IEEE Access.




