phishvision-zero-day-phishing-corpus
收藏资源简介:
OpticParse 是一个边缘原生计算机视觉网页抓取与威胁情报引擎,该数据集包含从150个持续运行的提取管道中收集的结构化数据,这些管道全天候在边缘运行,无需任何CSS选择器。数据集包含四个子集:主数据集(master_150_templates_dataset.csv)为多区域目录,包含调度信息和类别;威胁情报数据集(threat_intel_dataset.csv)包含已验证的钓鱼套件IOC、品牌仿冒URL和加密货币盗取器;电商与零售数据集(opticparse_ecommerce_and_retail.csv)提供亚马逊、Shopify和暗店实时定价遥测;B2B与金融数据集(opticparse_b2b_and_finance.csv)包含高管招聘信号、YC创始人信息和SEC 10-K文件。数据集规模为10K到100K条记录,单语种英语,适用于特征提取、文本检索、零样本分类等任务,特别适用于文档检索、网络安全、电商定价、B2B销售线索等场景。
OpticParse is an edge-native computer vision web scraping and threat intelligence engine. The dataset contains structured data collected from 150 continuously running extraction pipelines that operate 24/7 at the edge without any CSS selectors. The dataset consists of four subsets: the master dataset (master_150_templates_dataset.csv) is a multi-region catalog containing scheduling information and categories; the threat intelligence dataset (threat_intel_dataset.csv) contains verified phishing kit IOCs, brand impersonation URLs, and cryptocurrency stealers; the e-commerce and retail dataset (opticparse_ecommerce_and_retail.csv) provides real-time pricing telemetry from Amazon, Shopify, and dark stores; the B2B and finance dataset (opticparse_b2b_and_finance.csv) contains executive recruitment signals, YC founder information, and SEC 10-K filings. The dataset size ranges from 10K to 100K records, monolingual English, and is suitable for tasks such as feature extraction, text retrieval, zero-shot classification, especially for document retrieval, cybersecurity, e-commerce pricing, B2B sales leads, etc.
数据集概述
OpticParse 是一个包含150种模板的多行业智能语料库,由边缘原生计算机视觉网页抓取与威胁情报引擎生成。该语料库数据来自150条持续运行的提取管道,以零CSS选择器方式在边缘端全天候运行。
包含的数据文件
| 文件名 | 内容描述 |
|---|---|
master_150_templates_dataset.csv |
完整的多区域目录,包含调度信息和类别 |
threat_intel_dataset.csv |
已验证的网络钓鱼工具包IOC、品牌仿冒URL及加密货币盗取器 |
opticparse_ecommerce_and_retail.csv |
来自Amazon、Shopify和暗店(dark-store)的实时定价遥测数据 |
opticparse_b2b_and_finance.csv |
高管招聘信号、YC创始人和SEC 10-K申报数据流 |
数据集属性
- 语言:英语(单语)
- 许可证:CC0-1.0
- 数据规模:10K < n < 100K
- 标注方式:机器生成
- 任务类别:特征提取、文本检索、零样本分类
该数据集主要面向网络安全威胁情报、电子商务定价监控以及B2B金融数据挖掘等应用场景,支持多模态视觉与文本检索的联合分析。





