population-india-uttar-pradesh-cohort-100
收藏资源简介:
HipAAsynth数据集是一个用于验证的合成数据集,由HipAAsynth生成,旨在模拟真实世界的变异性以评估医疗系统在部署条件下的表现。该数据集代表了一个用于测试和基准测试的受控队列,模拟了患者群体、人口统计分布和共病模式的呈现。数据集设计用于可重复的评估和跨系统的一致比较。数据集包含29个结构化CSV列,涵盖患者标识符、人口统计信息(如年龄、性别、地区类型)、临床诊断(如主要和次要诊断)、医疗支付类型、糖尿病分类、疫苗接种状态、医疗访问权限等。数据集适用于表格分类和表格回归任务,特别适合医疗健康领域的评估和研究。
The HipAAsynth dataset is a synthetic validation dataset generated by HipAAsynth, which is designed to simulate real-world variability for evaluating the performance of medical systems under operational deployment conditions. This dataset constitutes a controlled cohort for testing and benchmarking, replicating the profiles of patient populations, demographic distributions, and comorbidity patterns. The dataset is engineered to support reproducible evaluation and consistent cross-system comparison. It contains 29 structured CSV columns covering patient identifiers, demographic information (e.g., age, gender, regional type), clinical diagnoses (e.g., primary and secondary diagnoses), medical payment types, diabetes classification, vaccination status, medical access permissions, and other related fields. The dataset is suitable for tabular classification and tabular regression tasks, and is particularly well-suited for healthcare domain evaluation and research.
数据集概述
基本元数据
- 数据集名称: HipAAsynth Dataset
- 托管地址: https://huggingface.co/datasets/HipAAsynth/population-india-uttar-pradesh-cohort-100
- 版本: 1.0.0
- 许可协议: cc-by-nc-4.0
- 数据规模类别: n<1K
- 任务类别: 表格分类、表格回归
- 标签: 医疗健康、电子健康记录、表格数据、评估、确定性
数据集描述
- 性质: 由HipAAsynth生成的验证工件,用于测试和基准评估。
- 目的: 模拟真实世界的变异性,以评估医疗系统在部署条件下的性能。
- 设计原则: 用于可重复评估和跨系统一致比较。
- 模拟内容: 模拟患者群体、人口统计分布和共病模式在不同条件下的呈现。
数据结构
- 格式: 29列结构化CSV文件。
- 数据文件: synthetic_india_uttar_pradesh_sample_100_seed93.csv
- 数据分割: 训练集
字段说明
| 字段名 | 描述 |
|---|---|
| patient_id | 唯一患者标识符 |
| data_type | 数据集类型标识符 |
| country | 来源国家 |
| state | 印度邦 |
| region_type | 城市或农村分类 |
| age | 年龄(岁) |
| sex | 生理性别 |
| language_region | 区域语言组 |
| primary_diagnosis | 主要临床诊断 |
| secondary_diagnosis | 次要临床诊断 |
| payer | 支付方类型(公共/私人/自费) |
| diabetes_type | 糖尿病分类(1型/2型/无) |
| hba1c | 糖化血红蛋白值 |
| bmi | 身体质量指数 |
| anemia | 贫血标志 |
| tb_exposure | 结核病暴露标志 |
| vaccination_covid19 | COVID-19疫苗接种状态 |
| vaccination_polio | 脊髓灰质炎疫苗接种状态 |
| vaccination_bcg | 卡介苗接种状态 |
| healthcare_access_primary | 初级医疗保健可及性 |
| healthcare_access_secondary | 二级医疗保健可及性 |
| healthcare_access_tertiary | 三级医疗保健可及性 |
| anchor_hash | 用于可重复性的SHA-256生成锚点 |
| facility_country | 医疗机构所在国家 |
| data_nature | 数据性质(合成) |
| generated_by | 生成引擎标识符 |
| organization | 发起组织 |
| license_engine | 引擎许可类型 |
| license_data | 数据许可类型 |




