遇见数据集

CreditTransAct: A Profile-Driven Dataset for Scalable Credit Card Fraud Detection

收藏
Mendeley Data2026-05-21 收录
官方服务:

资源简介:

A large-scale synthetic dataset of 15 million credit card transactions is generated for fraud detection research. It is divided into four segments: A_Established (40%), B_Regular (30%), C_New (20%), and D_Guest (10%), with segment-wise fraud rates of 2.5%, 4.0%, 7.0%, and 12.0% respectively. Overall, 95.10% of transactions are legitimate and 4.90% are fraudulent. Segments A, B, and C are based on persistent customer profiles that include behavioral attributes such as baseline spending, usual merchant category, account age, and credential change history. Segment D represents anonymous transactions generated from global distributions without historical customer context. Each record includes 33 features covering transaction amount, geographic behavior, device and network signals, authentication details, and behavioral patterns. Fraud cases are generated using four compound signal clusters: Card-Not-Present, Bot/Card Testing, Account Takeover, and Geo-Velocity. Controlled label noise is introduced in borderline cases to simulate real-world uncertainty. The dataset is generated using a reproducible and memory-efficient pipeline built with Python, NumPy, Pandas, and PyArrow. It is provided in Snappy-compressed Parquet format for efficient storage and in CSV format for easy accessibility. This dataset is suitable for understanding and assessing fraud behavior in realistic financial transaction settings.

本数据集为面向欺诈检测研究构建的大规模合成信用卡交易数据集,总计包含1500万条交易记录。 数据集被划分为四个分段:A_已留存客户(A_Established,占比40%)、B_常规客户(B_Regular,占比30%)、C_新客户(C_New,占比20%)以及D_访客用户(D_Guest,占比10%),各分段的欺诈率分别为2.5%、4.0%、7.0%与12.0%。整体而言,95.10%的交易为合法交易,仅4.90%为欺诈交易。 分段A、B与C均基于持久化客户画像构建,涵盖基础消费基准、常用商户类别、账户存续时长以及凭证变更历史等行为属性;分段D则为基于全局分布生成的匿名交易,不携带历史客户上下文信息。 每条交易记录包含33个特征字段,覆盖交易金额、地理行为、设备与网络信号、认证详情以及行为模式等多个维度。 欺诈案例通过四类复合信号簇生成:非现场刷卡(Card-Not-Present)、机器人/卡片测试(Bot/Card Testing)、账户接管(Account Takeover)以及地域流速异常(Geo-Velocity)。 研究人员在边界样本中引入可控的标签噪声,以模拟真实场景中的不确定性。 本数据集通过基于Python、NumPy、Pandas与PyArrow构建的可复现且内存高效的流水线生成。数据集以Snappy压缩的Parquet格式存储以实现高效存储,同时提供CSV格式以保障易用性。 本数据集适用于在真实金融交易场景中理解与评估欺诈行为模式。

创建时间:
2026-05-09
二维码
社区交流群
二维码
科研交流群
商业服务