Amharic text dataset extracted from memes for hate speech detection or classification: Amharic Language Hate Speech Detection System from Facebook Image Post Using Deep Learning System
收藏资源简介:
the dataset is collected from social media such as facebook and telegram. the dataset is further processed. the collection are D1_org: this dataset is neither stemed nor stopword are remove: D1_sf: in this dataset stopwords are removed but not stemmed and in D3_stemed datset is stemmed and stopwords are removed. stemming is done using hornmorpho developed by Michael Gesser( available at https://github.com/hltdi/HornMorpho) all datasets are normalized and free from noise such as punctuation marks and emojs.
本数据集采集自Facebook、Telegram等社交媒体平台。本数据集已完成进一步处理,其包含三个子数据集:D1_org:该数据集未进行词干提取(stemming),也未移除停用词(stopword);D1_sf:该数据集已移除停用词但未进行词干提取;D3_stemed:该数据集同时完成了词干提取与停用词移除。本次词干提取采用Michael Gesser开发的HornMorpho工具,可从https://github.com/hltdi/HornMorpho获取。所有数据集均已完成归一化处理,且已清除标点符号、表情符号等噪声数据。



