遇见数据集

Image dataset to train a deep learning model to decode Leetspeak obfuscated characters

收藏
Zenodo2022-03-21 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The dataset contains an image database (18,981 images) that could be used to train a deep learning model to accurately detect characters. We have successfully used it to create a model that identifies characters encoded using LeetSpeak. The original dataset can be found in the Mondragon Unibertsitatea Repository -- https://gitlab.danz.eus/datasharing/ski4spam The training dataset consists of: - Alphabetic letters (a-z) written using different fonts and styles (regular, cursive, bold, cursive+bold) - Handwritten letters: English handwriting from the Chars74k dataset [2] which is available at http://www.ee.surrey.ac.uk/CVSSP/demos/chars74k/.

本数据集包含一个图像数据库(共18981张图像),可用于训练深度学习模型以实现高精度字符检测。本团队已基于该数据集成功搭建可识别LeetSpeak编码字符的模型。原始数据集可于蒙德龙大学(Mondragon Unibertsitatea)仓库获取,链接为:https://gitlab.danz.eus/datasharing/ski4spam。 训练数据集包含以下内容: - 采用多种字体与样式(常规体、草书体、粗体、草书加粗体)书写的英文字母(a-z) - 手写字符:源自Chars74k数据集(Chars74k dataset)[2]的英文手写字符,该数据集可通过http://www.ee.surrey.ac.uk/CVSSP/demos/chars74k/ 公开获取。

提供机构:
Zenodo
创建时间:
2022-03-21
二维码
社区交流群
二维码
科研交流群
商业服务