Dataset of 1000 Images of Malicious and Benign QR codes 2025
收藏资源简介:
This comprehensive dataset comprises 1,000 high-quality QR code images, methodically organized into two distinct categories: 500 QR codes containing verified malicious URLs with documented phishing or malware distribution history 500 QR codes containing legitimate (benign) URLs from trusted domains The dataset addresses a critical gap in contemporary cybersecurity research materials, as most publicly available QR code datasets are significantly outdated and fail to reflect current cyber threat landscapes. Each QR code has been generated at a standardized resolution of 330x330 pixels to facilitate consistent processing across different machine learning frameworks. This collection is specifically curated to support: Binary classification algorithm development and evaluation Computer vision-based phishing detection research QR code security analysis and threat modeling Transfer learning applications in cybersecurity All malicious URLs have been obtained from multiple threat intelligence sources and scanning services, such as Virus Total, URLhaus, PhishTak etc, while benign URLs have been selected from Alexa top-ranked domains to ensure legitimacy. The dataset includes a diverse range of URL structures, encoding densities, and error correction levels to better represent real-world QR code variations. Researchers can utilize this dataset to train, validate, and test machine learning models for automated QR code threat detection, contributing to improved mobile security solutions in an era of increasing QR code adoption for financial transactions and authentication.
本综合数据集包含1000张高质量二维码 (QR Code) 图像,被系统划分为两个截然不同的类别: 500个二维码包含已验证的恶意URL,这些URL均带有被记录的网络钓鱼或恶意软件分发历史; 500个二维码包含合法(良性)URL,来源均为可信域名。 该数据集填补了当代网络安全研究资料中的关键空白——当前绝大多数公开可用的二维码数据集均已严重过时,无法反映当前的网络威胁态势。每张二维码均以330×330像素的标准分辨率生成,以确保在不同机器学习框架下均可实现一致的处理。 本数据集专为支持以下研究与开发工作而精心构建: 1. 二分类算法的开发与评估 2. 基于计算机视觉的钓鱼检测研究 3. 二维码安全分析与威胁建模 4. 网络安全领域的迁移学习应用 所有恶意URL均来源于多个威胁情报源及扫描服务,例如VirusTotal、URLhaus、PhishTak等;而良性URL则选自Alexa排名靠前的域名,以确保其合法性。数据集涵盖了多样化的URL结构、编码密度及纠错等级,以更贴近真实场景中的二维码变体。 研究人员可利用本数据集训练、验证并测试用于自动化二维码威胁检测的机器学习模型,助力在二维码日益广泛应用于金融交易与身份验证的当下,优化移动安全解决方案。




