Amharic text dataset extracted from memes for hate speech detection or classification
收藏资源简介:
the dataset is collected from social media such as facebook and telegram. the dataset is further processed. the collection are orginal_cleaned: this dataset is neither stemed nor stopword are remove: stopword_removed: in this dataset stopwords are removed but not stemmed and in stemed datset is stemmed and stopwords are removed. stemming is done using hornmorpho developed by Michael Gesser( available at https://github.com/hltdi/HornMorpho) all datasets are normalized and free from noise such as punctuation marks and emojs.
本数据集采集自Facebook、Telegram等社交媒体平台。 该数据集已完成进一步处理,共分为三类子集:其一为orginal_cleaned:此类数据集既未进行词干提取(stemming),也未移除停用词(stopword);其二为stopword_removed:此类数据集已移除停用词,但未进行词干提取;其三为词干提取版数据集,该类数据集既完成了词干提取,也移除了停用词。本次词干提取采用Michael Gesser开发的HornMorpho工具,其开源地址为https://github.com/hltdi/HornMorpho。 所有数据集均已完成归一化处理,且已清除标点符号、表情符号等噪声数据。




