Set of obfuscated spam dataset by using LeetSpeak transformations
收藏资源简介:
The usage of LeetSpeak and other text hiding tricks is often used by spammers in the distribution of unsolicited contents. To evaluate deobfuscation techniques and their impact on spam content classification, we preprocessed several popular public datasets to partially obfuscate the text. The datasets transformed are: YouTube Spam Collection [2, 3] which is available on https://www.dt.fee.unicamp.br/~tiago/youtubespamcollection/. a subset of YouTube Comments [4, 5] which is available on http://mlg.ucd.ie/yt/. CSDMC2010 which is available on http://csmining.org/index.php/spam-email-datasets-.html. TREC2007 which is available on https://plg.uwaterloo.ca/~gvcormac/treccorpus07/
垃圾邮件发送者常使用Leet语(LeetSpeak)及其他文本隐藏技巧来分发未经请求的内容。为了评估文本反混淆(deobfuscation)技术及其对垃圾邮件内容分类的影响,我们对多个热门公开数据集进行了预处理,对其中的文本执行部分混淆操作。本次经过转换处理的数据集如下:YouTube垃圾邮件数据集(YouTube Spam Collection)[2, 3],其公开获取地址为https://www.dt.fee.unicamp.br/~tiago/youtubespamcollection/;YouTube评论数据集子集(YouTube Comments)[4, 5],公开获取地址为http://mlg.ucd.ie/yt/;CSDMC2010数据集,公开获取地址为http://csmining.org/index.php/spam-email-datasets-.html;TREC2007数据集,公开获取地址为https://plg.uwaterloo.ca/~gvcormac/treccorpus07/



