遇见数据集

医疗保险欺诈检测数据集

收藏
海数据2026-03-14 收录
官方服务:

资源简介:

医疗保险欺诈检测数据集_Healthcare_Fraud_Detection_Dataset 数据来源:互联网公开数据 标签:医疗保险, 欺诈检测, 医疗服务, 医生信息, 费用分析, 机器学习, 数据挖掘, 风险评估 数据概述: 该数据集包含来自医疗保险理赔的数据,记录了医疗服务提供者的相关信息、医疗服务类型、费用以及欺诈标签。主要特征如下: 时间跨度:数据未明确标注时间范围,可视为静态数据集。 地理范围:数据未明确标注地理范围,但从数据字段推测可能来源于美国。 数据维度:数据集包含多个维度,包括医生信息(如NPI号码、姓名、执业类型、邮编等)、医疗服务信息(如HCPCS代码、服务描述)、费用信息(如总受益人数、总服务次数、平均申报费用、平均允许费用、平均医疗保险支付金额、平均标准化金额)以及欺诈标签(Label,0表示非欺诈,1表示欺诈)。 数据格式:CSV格式,文件名为label_data.csv,便于数据分析与处理。 来源信息:数据来源可能为公开的医疗保险理赔数据,已进行匿名化处理。 该数据集适合用于医疗保险欺诈检测、风险评估和医疗服务费用分析等领域。 数据用途概述: 该数据集具有广泛的应用潜力,特别适用于以下场景: 研究与分析:适用于医疗健康领域学术研究,如医疗欺诈行为模式分析、异常理赔识别、医疗费用影响因素分析等。 行业应用:为医疗保险公司、医疗服务提供商提供数据支持,用于风险管理、欺诈检测、理赔审核流程优化等。 决策支持:支持医疗保险行业的风险控制策略制定,辅助保险公司进行欺诈调查和预防。 教育和培训:作为医疗健康数据分析、机器学习、风险管理等课程的实训材料,帮助学生和研究人员理解医疗欺诈检测的实践应用。 此数据集特别适合用于探索医疗服务费用与欺诈行为之间的关系,帮助用户建立欺诈检测模型、优化风险控制措施。

Healthcare Fraud Detection Dataset Data Source: Publicly available data from the Internet Labels: Medical Insurance, Fraud Detection, Healthcare Services, Provider Information, Cost Analysis, Machine Learning, Data Mining, Risk Assessment Data Overview: This dataset contains data from medical insurance claims, recording relevant information of healthcare providers, types of medical services, costs, and fraud labels. The main features are as follows: Time Span: No explicit time range is specified for the data, so it can be regarded as a static dataset. Geographic Scope: No explicit geographic range is specified, but it is inferred to originate from the United States based on the data fields. Data Dimensions: The dataset includes multiple dimensions, covering provider information (e.g., National Provider Identifier (NPI) number, name, practice type, zip code, etc.), medical service information (e.g., Healthcare Common Procedure Coding System (HCPCS) code, service description), cost information (e.g., total number of beneficiaries, total number of services, average submitted charge, average allowed charge, average Medicare payment amount, average standardized amount), and fraud labels (Label, where 0 indicates non-fraudulent and 1 indicates fraudulent). Data Format: In CSV format, with the file name label_data.csv, facilitating data analysis and processing. Source Information: The data is likely sourced from publicly available medical insurance claim data and has been anonymized. This dataset is suitable for applications in fields such as medical insurance fraud detection, risk assessment, and healthcare service cost analysis. Data Usage Overview: This dataset has broad application potential and is particularly suitable for the following scenarios: Research and Analysis: Suitable for academic research in the healthcare field, such as analysis of medical fraud behavior patterns, abnormal claim identification, analysis of influencing factors of medical costs, etc. Industry Applications: Providing data support for medical insurance companies and healthcare service providers, used for risk management, fraud detection, optimization of claim review processes, etc. Decision Support: Supporting the formulation of risk control strategies in the medical insurance industry, assisting insurance companies in fraud investigation and prevention. Education and Training: Serving as practical training materials for courses such as healthcare data analysis, machine learning, and risk management, helping students and researchers understand the practical applications of medical fraud detection. This dataset is particularly suitable for exploring the relationship between healthcare service costs and fraudulent behaviors, helping users build fraud detection models and optimize risk control measures.

提供机构:
互联网公开数据
创建时间:
2026-03-11
搜集汇总
数据集介绍
医疗保险欺诈检测数据集 数据集图片
背景与挑战
背景概述
该数据集是一个用于医疗保险欺诈检测的公开数据集,包含医生信息、医疗服务类型、费用数据以及欺诈标签(0为非欺诈,1为欺诈),适用于机器学习模型构建和风险评估研究。数据以CSV格式提供,覆盖多维度特征如NPI号码、服务描述和平均费用,主要应用于医疗健康领域的欺诈行为分析、理赔优化和学术培训。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务