electricsheepasia/asia-ilo-how-temp-sex-eco-edu-nb-mean-weekly-hours-actually-worked-per-employed-per
收藏资源简介:
该数据集包含了亚洲32个国家从1997年至2025年间的359,636条观测数据,核心指标为平均每周实际工作时长(Mean weekly hours actually worked per employed person),按性别、经济活动和教育程度进行细分。数据来源于国际劳工组织(ILO)的ILOSTAT统计数据库,通过其REST API直接获取,并过滤为亚洲国家代码。数据集采用表格形式,包含多个字段:国家代码(ref_area)、国家名称(ref_area.label)、数据来源(source.label)、指标代码(indicator)、性别分类(sex)、经济活动和教育分类(classif1, classif2)、观测年份(time)、观测值(obs_value)以及数据状态标志(obs_status)等。该数据集适用于机器学习任务,如表格分类、回归分析和时间序列预测,可用于研究亚洲劳动力市场的工作时间趋势、性别差异、经济部门影响等。数据经过ILO harmonisation处理,确保符合国际劳工统计会议(ICLS)定义,并标注了来源以便追溯。数据集以Parquet格式发布,方便通过HuggingFace的datasets库直接加载使用。
This dataset contains 359,636 observations of Hours of work data across 32 Asia countries, spanning from 1997 to 2025. The primary indicator is Mean weekly hours actually worked per employed person by sex, economic activity and education (HOW_TEMP_SEX_ECO_EDU_NB). Data is sourced from the International Labour Organization (ILO) ILOSTAT database, retrieved via REST API and filtered to Asia ISO3 country codes. The dataset is structured in tabular format with columns including country code (ref_area), country name (ref_area.label), data source (source.label), indicator code (indicator), sex disaggregation (sex), economic activity and education classifications (classif1, classif2), observation year (time), observed value (obs_value), and observation status flags (obs_status). It is designed for machine learning tasks such as tabular classification, regression, and time-series forecasting, enabling analysis of labor market trends, gender disparities, and sectoral impacts in Asia. Data is harmonized using ICLS definitions and includes source traceability. The dataset is packaged as Parquet and accessible via HuggingFace Datasets for easy integration.




