遇见数据集

Real and Synthetic Income Datasets for Comparative Statistical and Machine Learning Analysis

收藏
Mendeley Data2026-09-08 收录
官方服务:

资源简介:

This dataset contains a real Adult Income dataset and an author-generated synthetic income dataset prepared for comparative statistical and machine learning analysis. The real dataset is based on the Adult dataset from the UCI Machine Learning Repository, while the synthetic dataset was developed as a Bangladesh-oriented tabular income dataset using demographic, educational, employment, working-hour, experience, and income-related variables. The datasets are provided to support the comparison of real and synthetic tabular income data in terms of their structure, distributions, statistical relationships, and usefulness for machine learning experiments. The accompanying analysis includes data inspection, preprocessing, descriptive statistics, correlation analysis, statistical testing, classification experiments, class-imbalance handling, and model evaluation. The machine learning analysis includes Logistic Regression, Decision Tree, Random Forest, and XGBoost, with additional experiments using class weighting and Synthetic Minority Over-sampling Technique (SMOTE) where applicable. Evaluation measures include accuracy, precision, recall, F1-score, ROC-AUC, and Balanced Error Rate (BER). The synthetic dataset is intended as a controlled data resource for methodological comparison and experimentation. It should not be interpreted as nationally representative or as observed income data from the Bangladeshi population. In particular, the relationships within the synthetic data may reflect the assumptions and construction process used to generate the data rather than naturally occurring socioeconomic relationships. The repository is intended to support transparency, reproducibility, and future reuse of the datasets and associated analysis materials. The real and synthetic datasets should be considered separately because their variables, target representations, and data-generation processes differ.

本数据集包含真实成人收入数据集与作者生成的合成收入数据集,专为对比统计分析与机器学习分析而制备。其中真实数据集基于UCI机器学习库(UCI Machine Learning Repository)的成人数据集(Adult dataset),而合成数据集则是一款面向孟加拉国的表格型收入数据集,采用人口统计、教育、就业、工作时长、工作经验及收入相关变量构建。 本数据集旨在支持真实与合成表格型收入数据在结构、分布、统计关联及机器学习实验可用性层面的对比研究。配套的分析流程涵盖数据检视、预处理、描述性统计、相关性分析、统计检验、分类实验、类别不平衡处理及模型评估。 本次机器学习分析涵盖逻辑回归(Logistic Regression)、决策树(Decision Tree)、随机森林(Random Forest)及XGBoost,针对适用场景额外开展了类别权重调整与合成少数类过采样技术(Synthetic Minority Over-sampling Technique, SMOTE)相关实验。评估指标包括准确率、精确率、召回率、F1值、ROC-AUC及平衡错误率(Balanced Error Rate, BER)。 本合成数据集旨在作为受控数据资源,用于方法论对比与实验研究。本数据集不可被解读为具有全国代表性的数据,亦不可视为孟加拉国人口的真实收入观测数据。特别地,合成数据内的关联关系可能反映的是数据生成所依据的假设与构建流程,而非自然存在的社会经济关联。 本数据集仓库旨在提升数据集及配套分析材料的透明度、可复现性与未来复用性。真实数据集与合成数据集应分开使用,因为二者的变量、目标变量表征及数据生成流程均存在差异。

创建时间:
2026-09-03
二维码
社区交流群
二维码
科研交流群
商业服务