遇见数据集

Phish-Iris Dataset: A Small Scale Multi-Class Phishing Web Page Screenshots Dataset

收藏
Mendeley Data2024-01-31 更新2024-06-26 收录
官方服务:

资源简介:

Phish-IRIS dataset is aimed for researchers to supply a ground truth dataset to evaluate their vision based multi-class anti-phishing studies. For this purpose, we supply a corpus involving unique screenshots of 15 (14+1) classes. Here, 1 class represents the "unknown" or "legitimate" samples while the rest of the 14 classes correspond to different highly phished brands. It is important to mention that, Phish-IRIS dataset aims to provide a benchmark dataset for only computer vision based anti-phishing studies. The dataset has been collected between March and May 2018. The phishing pages have been collected from Phishtank.com and OpenPhish.com while legitimate pages have been collected randomly. Our dataset involves 1313 training and 1539 testing samples. The directory structure of the dataset has been splitted as "train" and "val" folders which contain respective brand names and "other" category. Since the nature of the anti-phishing is based on discriminating legitimate web pages from the phished targets, we also provided a fairly larger set for legitimate samples (i.e. "other" category). Our first paper utilizing this dataset has been published with the title of Phish-IRIS: A New Approach for Vision Based Brand Prediction of Phishing Web Pages via Compact Visual Descriptors. The official home page of the dataset is https://web.cs.hacettepe.edu.tr/~selman/phish-iris-dataset/ We believe that, our dataset will be beneficial for the researchers who are interested in vision based anti phishing .

Phish-IRIS数据集旨在为研究人员提供带有真实标注的基准数据集,用于评估基于视觉的多分类反钓鱼相关研究。本数据集涵盖15(14+1)个类别的唯一截图语料库:其中1个类别代表"未知"或"合法"样本,其余14个类别对应不同的热门仿冒品牌。需特别说明的是,Phish-IRIS数据集仅面向基于计算机视觉的反钓鱼研究打造。该数据集采集于2018年3月至5月期间,钓鱼页面采集自Phishtank.com与OpenPhish.com,合法页面则通过随机方式获取。本数据集共包含1313个训练样本与1539个测试样本。数据集的目录结构分为"train"与"val"两个文件夹,分别收纳各对应品牌类别与"other"分类。鉴于反钓鱼任务的核心在于区分合法网页与仿冒目标,我们为合法样本(即"other"分类)设置了相对更大的样本规模。首个使用本数据集的研究论文以《Phish-IRIS:基于紧凑视觉描述符的钓鱼网页视觉品牌识别新方法》为题发表。该数据集的官方主页为https://web.cs.hacettepe.edu.tr/~selman/phish-iris-dataset/。我们相信,本数据集将为关注基于视觉的反钓鱼研究的科研人员提供切实助力。

创建时间:
2024-01-31
搜集汇总
数据集介绍
Phish-Iris Dataset: A Small Scale Multi-Class Phishing Web Page Screenshots Dataset 数据集图片
背景与挑战
背景概述
Phish-Iris数据集是一个小规模多类别钓鱼网页截图数据集,专为基于视觉的反钓鱼研究设计。它包含15个类别(14个钓鱼品牌和1个合法/未知类别),共2852个样本(1313个训练样本和1539个测试样本),数据收集于2018年,来源包括公开钓鱼网站和随机合法页面,旨在为计算机视觉方法提供基准评估。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务