遇见数据集

creditcard Dataset

收藏
Figshare2025-06-09 更新2026-04-08 收录
官方服务:

资源简介:

<b>Title:</b> Credit Card Transactions Dataset for Fraud Detection (Used in: A Hybrid Anomaly Detection Framework Combining Supervised and Unsupervised Learning)<b>Description:</b>This dataset, commonly known as <i>creditcard.csv</i>, contains anonymized credit card transactions made by European cardholders in September 2013. It includes 284,807 transactions, with 492 labeled as fraudulent. Due to confidentiality constraints, features have been transformed using PCA, except for 'Time' and 'Amount'.This dataset was used in the research article titled <i>"A Hybrid Anomaly Detection Framework Combining Supervised and Unsupervised Learning for Credit Card Fraud Detection"</i>. The study proposes an ensemble model integrating techniques such as Autoencoders, Isolation Forest, Local Outlier Factor, and supervised classifiers including XGBoost and Random Forest, aiming to improve the detection of rare fraudulent patterns while maintaining efficiency and scalability.<b>Key Features:</b>30 numerical input features (V1–V28, Time, Amount)Class label indicating fraud (1) or normal (0)Imbalanced class distribution typical in real-world fraud detection<b>Use Case:</b><br>Ideal for benchmarking and evaluating anomaly detection and classification algorithms in highly imbalanced data scenarios.<b>Source:</b><br>Originally published by the Machine Learning Group at Université Libre de Bruxelles.<br>https://www.kaggle.com/mlg-ulb/creditcardfraud<b>License:</b><br>This dataset is distributed for academic and research purposes only. Please cite the original source when using the dataset.

<b>标题:</b>用于欺诈检测的信用卡交易数据集(应用于:结合监督与无监督学习的混合异常检测框架)<b>描述:</b>本数据集通常以<i>creditcard.csv</i>为名,收录了2013年9月欧洲持卡人完成的匿名化信用卡交易记录。数据集总计包含284807笔交易,其中492笔被标记为欺诈交易。出于保密约束,除'Time'与'Amount'字段外,其余特征均经过主成分分析(PCA)转换。本数据集曾被用于题为<i>《用于信用卡欺诈检测的结合监督与无监督学习的混合异常检测框架》</i>的研究论文。该研究提出了一种集成模型,整合了自编码器(Autoencoders)、孤立森林(Isolation Forest)、局部离群因子(Local Outlier Factor)等无监督技术,以及极端梯度提升树(XGBoost)、随机森林(Random Forest)等监督分类器,旨在在保证运算效率与可扩展性的同时,提升对罕见欺诈模式的检测能力。<b>关键特征:</b>30个数值型输入特征(V1–V28、Time、Amount);用于标识交易欺诈状态的类别标签(1代表欺诈交易,0代表正常交易);具备现实欺诈检测场景中典型的非平衡类别分布。<b>应用场景:</b><br>非常适合作为基准数据集,用于评估高度非平衡数据场景下的异常检测与分类算法性能。<b>数据来源:</b><br>最初由布鲁塞尔自由大学机器学习小组(Machine Learning Group at Université Libre de Bruxelles)发布。<br>https://www.kaggle.com/mlg-ulb/creditcardfraud<b>使用许可:</b><br>本数据集仅可用于学术与研究用途,使用该数据集时请引用原始数据源。

创建时间:
2025-06-09
二维码
社区交流群
二维码
科研交流群
商业服务