遇见数据集

fraud_oracle.csv

收藏
DataCite Commons2024-01-13 更新2024-08-19 收录
官方服务:

资源简介:

This study leverages advanced methods and algorithms to create an automated insurance fraud detection system, using real insurance fraud data. The system consists of four phases: data resampling (Over, Under, and hybrid), feature selection (Filtering, Wrapping, and Embedding), binary classification (Bagging and Boosting), and explanatory model analysis (Shapley Additive Explanations, Break-down plots, and variable-importance Measures). Results show that not all resampling techniques improve algorithm performance, but all feature selection methods do. Notably, the Boosting algorithm, incorporating the Neighborhood Cleaning Rule for resampling and Tree-based feature selection, excels in detecting insurance claim fraud.

本研究借助先进方法与算法,基于真实保险欺诈数据搭建了自动化保险欺诈检测系统(automated insurance fraud detection system)。该系统包含四大研究环节:其一为数据重采样(Data Resampling),涵盖过采样(Over-sampling)、欠采样(Under-sampling)与混合采样(Hybrid sampling)三种方式;其二为特征选择(Feature Selection),包含过滤法(Filtering)、包装法(Wrapping)与嵌入法(Embedding)三类方法;其三为二分类任务(Binary Classification),采用装袋法(Bagging)与提升法(Boosting)两种集成学习算法;其四为可解释性模型分析(Explanatory Model Analysis),涵盖夏普利可加解释(Shapley Additive Explanations)、分解图(Break-down plots)与变量重要性度量(Variable-importance Measures)三种分析手段。研究结果表明,并非所有重采样技术均可提升算法性能,但所有特征选择方法均能有效优化模型表现。值得关注的是,结合邻域清洁规则(Neighborhood Cleaning Rule)开展重采样并采用基于树的特征选择(Tree-based feature selection)的提升法(Boosting)算法,在保险索赔欺诈检测任务中表现最为优异。

提供机构:
figshare
创建时间:
2024-01-13
搜集汇总
数据集介绍
fraud_oracle.csv 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务