遇见数据集

Sample-weighting example: Population proportions.

收藏
NIAID Data Ecosystem2026-05-01 收录
官方服务:

资源简介:

In most countries, a government agency or collaborating organization gathers information on occupational accidents. Comparisons based on a single factor such as autonomous community, activity sector or others, often leads to contradictory conclusions. The use of this information for comparison is not immediate because the different characteristics considered give place to different possible comparisons. The elaboration of a single baseline for each set of characteristics is addressed. The method proposed comes from the data available in Spain but could be applied to other cases. The method consists of: (1) selecting factors–those selected are age, sex, autonomous community and activity; (2) the generation of a synthetic population based on data from a survey and general proportions by applying the Optimal Representative Sample Weighting (rsw); and (3) the prediction of the accidents ratio for each set of characteristic by using a XGBoost decision trees ensemble. The results confirm the appropriateness of the method.

在多数国家,政府机构或合作组织会收集职业事故相关信息。仅基于单一维度(如自治区、行业门类等)开展的对比分析,往往会得出相互矛盾的结论。由于需考量的特征维度各异,可实现的对比路径也各不相同,因此难以直接利用该类信息开展对比工作。本研究聚焦于为每一组特征维度构建统一基准的问题。所提出的方法基于西班牙现有数据构建,但亦可推广至其他场景。该方法包含以下步骤:(1) 特征选取:本次选取的特征包括年龄、性别、自治区以及行业活动类型;(2) 借助最优代表性样本加权(Optimal Representative Sample Weighting, rsw)方法,基于调研数据与总体比例生成合成人口群体;(3) 采用XGBoost决策树集成模型,预测每一组特征维度对应的事故发生率。实验结果验证了该方法的适用性。

创建时间:
2023-11-22
二维码
社区交流群
二维码
科研交流群
商业服务