遇见数据集

SendGuard900K: Massive Dataset for Metadata-Based Quality Assurance of Email Marketing

收藏
Zenodo2025-12-31 更新2026-05-26 收录
官方服务:

资源简介:

SendGuard900K: Massive Dataset for Metadata-Based Quality Assurance of Email Marketing This work was supported in part by the Project ‘‘SendGuard—Improving the Security of Recipients Innovative Tool Based on Machine Learning Technology and Artificial Intelligence to Fight with the Problem of Spam and Phishing in Marketing and Transactional E-Mail Messages" co-financed from the Funds of European Regional Development Fund under Project POIR.01.01.01-00-0202/19-02. Overview We introduce SendGuard900K, a large-scale real-world dataset designed to support research on quality assurance, security, and performance analysis of email marketing campaigns using metadata and aggregated behavioral signals. The dataset originates from the SendGuard project, whose goal is to develop a significantly improved AI-driven service for analyzing email marketing and transactional messages. The project addresses one of the most pressing global challenges in digital communication: ensuring that legitimate, personalized messages reach their recipients effectively in an ecosystem dominated by spam, phishing, and abuse. SendGuard900K contains rich campaign-level metadata collected from production systems between 2022 and 2024, spanning campaigns created by a large number of independent companies across multiple industries and geographies. The dataset is designed to enable reproducible research on email quality, deliverability, engagement, and abuse detection—without exposing message content or personally identifiable information. Files This dataset includes two CSV files: SendGuard900K.csv — The main file containing all records in the dataset. SendGuard900K-preview.csv — A sample of 100 rows extracted from the main file, provided solely for the purpose of enabling Zenodo’s preview functionality. Data summary The detail technical info about data including columns names, types and inteprpretation is given in the Technical Notes of this record. Time period: 2022–2024 Granularity: Campaign-level Scale: ~900,000 email campaigns Creators: Campaigns generated by a large number (over 4000) of independent companies Data types: Integer metrics, categorical descriptors, timestamps Privacy: No personal data, no message bodies, no recipient-level records Potential Research Use-cases The following examples illustrate some potential research directions enabled by the dataset; however, they are not exhaustive, and the actual scope of possible analyses and applications can easily extend beyond the cases listed below. 1. Deliverability and Engagement Prediction Task: Regression or classificationPotential targets: campaign_unique_opens_count campaign_unique_clicks_count derived open rate or click-through rate time-windowed engagement metrics (e.g. _30, _60, _1440) 2. Phishing Risk Modeling Task: Classification or risk scoring (regression)Potential targets: campaign_phishing_count campaign_complaint_count campaign_moderation campaign_moderation_rejected_count 3. Spam and Abuse Detection Task: Binary or multiclass classificationPotential targets: campaign_moderation campaign_moderation_rejected_count campaign_complaint_count campaign_phishing_count

SendGuard900K:用于电子邮件营销元数据质量保障的大规模数据集 本研究部分受名为"SendGuard——基于机器学习与人工智能技术开发面向收件人的创新工具,以应对营销与交易类电子邮件中的垃圾邮件与网络钓鱼问题"的项目资助,该项目由欧洲区域发展基金(European Regional Development Fund)通过POIR.01.01.01-00-0202/19-02号项目共同出资。 ## 概述 本研究提出SendGuard900K,这是一个大规模真实世界数据集,旨在支持基于元数据与聚合行为信号开展电子邮件营销活动的质量保障、安全防护与性能分析相关研究。 本数据集源自SendGuard项目,该项目旨在开发经过显著优化的人工智能驱动服务,用于分析电子邮件营销与交易类信息。本项目聚焦数字通信领域最紧迫的全球挑战之一:在垃圾邮件、网络钓鱼与滥用行为泛滥的生态环境中,确保合法的个性化邮件能够有效送达收件人。 SendGuard900K包含2022年至2024年间从生产系统中采集的丰富的活动级(Campaign-level)元数据,涵盖了来自多个行业、不同地区的大量独立企业所创建的营销活动。 本数据集旨在支持可复现的相关研究,涵盖邮件质量、送达率、用户参与度与滥用行为检测等方向,且不会泄露邮件内容或个人可识别信息。 ## 数据集文件 本数据集包含两个CSV文件: 1. SendGuard900K.csv:主数据集文件,包含本数据集全部记录。 2. SendGuard900K-preview.csv:从主文件中提取的100行样本,仅用于启用Zenodo平台的预览功能。 ## 数据摘要 本数据集的详细技术信息(包括字段名称、数据类型与含义阐释)已在本记录的技术说明中给出。 - 数据时间范围:2022年至2024年 - 数据粒度:活动级 - 数据规模:约90万个电子邮件营销活动 - 数据来源:由超过4000家独立企业创建的营销活动 - 数据类型:整数型指标、分类描述符、时间戳 - 隐私保护:不包含个人数据、邮件正文或收件人级别的记录 ## 潜在研究应用场景 以下示例列举了本数据集可支持的部分研究方向,但并非穷举,实际可开展的分析与应用范围远超下述案例。 1. 送达率与用户参与度预测 任务类型:回归或分类任务 潜在预测目标: - 活动唯一打开次数(campaign_unique_opens_count) - 活动唯一点击次数(campaign_unique_clicks_count) - 衍生打开率或点击率 - 时间窗口化的用户参与度指标(例如_30、_60、_1440) 2. 网络钓鱼风险建模 任务类型:分类或风险评分(回归)任务 潜在预测目标: - 活动钓鱼邮件计数(campaign_phishing_count) - 活动投诉计数(campaign_complaint_count) - 活动审核状态(campaign_moderation) - 活动审核拒绝次数(campaign_moderation_rejected_count) 3. 垃圾邮件与滥用行为检测 任务类型:二分类或多分类任务 潜在预测目标: - 活动审核状态(campaign_moderation) - 活动审核拒绝次数(campaign_moderation_rejected_count) - 活动投诉计数(campaign_complaint_count) - 活动钓鱼邮件计数(campaign_phishing_count)

提供机构:
Zenodo
创建时间:
2025-12-31
二维码
社区交流群
二维码
科研交流群
商业服务