遇见数据集

Syn-SWIFT: Synthetic SWIFT Transaction Dataset for Federated Fraud Detection

收藏
Zenodo2026-04-30 更新2026-05-26 收录
官方服务:

资源简介:

The inaccessibility of financial data constitutes a fundamental barrier to the development of robust AI models for frauddetection. Strict privacy regulations isolate repositories of data within an organization that cannot be accessed or shared by other departments, usually termed data silos, preventing institutions from pooling data for centralized training. This paper introduces Syn-SWIFT, a methodological framework for generating high-fidelity, synthetic, and labeled SWIFT message data. The framework is not a black-box generator; it is an explainable, hybrid system built on rigorously defined ground-truth constraints derived from expert domain knowledge, from 12-character BIC logic to incoming/outgoing message simulation to ensure all generated messages are semantically and structurally valid. The framework’s efficacy is demonstrated by generating four distinct datasets (MT101, MT103, MT110, MT202), totaling 400,000 transactions, segregated into 20 bank-specific silos. A modular pipeline incorporating fraud-pattern mutation, metadata synthesis, and raw message reconstruction ensures semantic correctness while maintaining operational variability. Each dataset is injected with a comprehensive lexicon of 40 unique, forensically sound fraud types, enabling complex multi-class classification. Experimental validation demonstrates the system’s effectiveness in producing coherent metadata–message alignment, stable reconstruction performance, and consistent fraud-type distribution across message types. Key challenges, including semantic drift, parser consistency, and multi-bank message fidelity, were addressed through rule-driven header enforcement and structured field regeneration. This siloed dataset is explicitly designed to serve as a new, public benchmark to accelerate research in Federated Learning (FL) for privacy-preserving financial crime detection.

金融数据的可及性不足,乃是构建鲁棒性欺诈检测人工智能(AI)模型的根本性阻碍。严格的隐私监管条例使得各机构内部的数据仓库相互隔离,其他部门无法访问或共享这些数据,这类数据隔离场景通常被称为数据孤岛(data silos),这极大阻碍了金融机构汇集数据以开展集中式模型训练。本文提出Syn-SWIFT框架,一种用于生成高保真、合成且带标注的环球银行金融电信协会(SWIFT)报文数据的方法论框架。该框架并非黑箱生成器,而是一套可解释的混合系统,其构建基于从专家领域知识中严格定义的真实约束条件——从12位银行识别码(BIC, Bank Identifier Code)逻辑,到传入/传出报文模拟,以确保所有生成的报文在语义与结构上均合法有效。本研究通过生成四类各具特色的数据集(MT101、MT103、MT110、MT202)验证了该框架的有效性,总计生成40万笔交易,并将其划分为20个银行专属数据孤岛。该框架采用模块化流水线设计,集成了欺诈模式变异、元数据合成与原始报文重构三大模块,在保证语义正确性的同时保留了业务操作的可变性。每个数据集均注入了涵盖40种独特且符合司法取证标准的欺诈类型的完整词库,可支持复杂的多分类任务。实验验证结果表明,该系统能够生成连贯的元数据-报文对齐结果,具备稳定的重构性能,且在不同报文类型间保持一致的欺诈类型分布。研究团队通过规则驱动的报文头强制规范与结构化字段重构,解决了语义漂移、解析器一致性以及多银行报文保真度等关键挑战。这款带有数据孤岛特性的数据集被专门设计为全新的公开基准测试集,以加速面向隐私保护型金融犯罪检测的联邦学习(Federated Learning, FL)相关研究。

提供机构:
Zenodo
创建时间:
2026-04-28
二维码
社区交流群
二维码
科研交流群
商业服务