Synthetic Stroke Prediction Dataset
收藏DataCite Commons2025-05-02 更新2025-05-17 收录
下载链接:
https://data.mendeley.com/datasets/s2nh6fm925
下载链接
链接失效反馈官方服务:
资源简介:
This dataset is a synthetic version inspired by the original "Stroke Prediction Dataset" on Kaggle. It contains anonymized, artificially generated data intended for research and model training on healthcare-related stroke prediction. The dataset generated using GPT-4o contains 50,000 records and 12 features. The target variable is stroke, a binary classification where 1 represents stroke occurrence and 0 represents no stroke. The dataset includes both numerical and categorical features, requiring preprocessing steps before analysis. A small portion of the entries includes intentionally introduced missing values to allow users to practice various data preprocessing techniques such as imputation, missing data analysis, and cleaning.
The dataset is suitable for educational and research purposes, particularly in machine learning tasks related to classification, healthcare analytics, and data cleaning.
No real-world patient information was used in creating this dataset.
本数据集是受Kaggle平台上原始《中风预测数据集》启发而构建的合成版本。它包含匿名化的人工生成数据,旨在支持医疗相关中风预测的研究与模型训练。该数据集通过GPT-4o生成,包含50,000条记录和12个特征。目标变量为“中风(stroke)”,属于二元分类问题——1代表中风发生,0代表未发生中风。数据集涵盖数值型与分类型特征,分析前需进行预处理。部分条目包含故意引入的缺失值,供用户练习插补(imputation)、缺失数据分析及数据清洗等多种预处理技术。
该数据集适用于教育与研究场景,尤其适合分类、医疗分析及数据清洗相关的机器学习任务。创建本数据集未使用任何真实患者信息。
提供机构:
Mendeley Data
创建时间:
2025-05-02



