遇见数据集

electricsheepasia/asia-ilo-eip-dwap-sex-edu-rt-inactivity-rate-by-sex-and-education

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - asia - ilostat - other-measures-of-labour-underutilization - ilo - labour - employment pretty_name: "Inactivity rate by sex and education (%) | Asia (ILOSTAT)" --- # Inactivity rate by sex and education (%) | Asia (ILOSTAT) 🌏 **17,694 observations** · **39 Asia countries** · **1970–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-17,694-blue) ![countries](https://img.shields.io/badge/countries-39-green) ![years](https://img.shields.io/badge/years-1970–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **17,694 observations** of `Other measures of labour underutilization` data across **39 Asia countries**, spanning **1970–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_DWAP_SEX_EDU_RT) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Other measures of labour underutilization ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EIP_DWAP_SEX_EDU_RT` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 39 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `IDN` | 1,245 | 1990 | 2023 | | `PSE` | 1,102 | 2000 | 2025 | | `CYP` | 1,069 | 1999 | 2024 | | `IRN` | 1,032 | 2005 | 2024 | | `KOR` | 996 | 2000 | 2025 | | `KHM` | 896 | 1996 | 2023 | | `TUR` | 882 | 2000 | 2024 | | `MNG` | 779 | 2003 | 2024 | | `PAK` | 732 | 2005 | 2025 | | `VNM` | 721 | 2010 | 2024 | | `THA` | 713 | 2000 | 2024 | | `ARM` | 692 | 2001 | 2023 | | `GEO` | 691 | 2009 | 2024 | | `IND` | 624 | 1994 | 2025 | | `ISR` | 624 | 2012 | 2024 | | ... | _24 more countries_ | | | ## Indicators (sample) - `EIP_DWAP_SEX_EDU_RT` — Inactivity rate by sex and education (%) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EIP_DWAP_SEX_EDU_RT` | | `indicator.label` | `string` | Indicator name in English | `Inactivity rate by sex and education (%)` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `50.27` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `B` | | `obs_status.label` | `string` | — | `Break in series` | | `note_classif` | `string` | — | `C3:2620` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (4 unique values): `SEX_T`, `SEX_M`, `SEX_F`, `SEX_O` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-eip-dwap-sex-edu-rt-inactivity-rate-by-sex-and-education") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EIP_DWAP_SEX_EDU_RT"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EIP_DWAP_SEX_EDU_RT") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EIP_DWAP_SEX_EDU_RT"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_eip_dwap_sex_edu_rt_inactivity_rate_by_sex_and_education_2025, title = {Inactivity rate by sex and education (%) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_DWAP_SEX_EDU_RT}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-eip-dwap-sex-edu-rt-inactivity-rate-by-sex-and-education}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_DWAP_SEX_EDU_RT_

This dataset contains 17,694 observations of Inactivity rate by sex and education (%) data across 39 Asia countries, spanning from 1970 to 2025. It is sourced from the ILOSTAT database of the International Labour Organization (ILO), extracted via API and filtered to Asian ISO3 country codes. The data is structured in tabular format, including columns such as country code, country name, source, indicator code, indicator name, sex disaggregation, education classification, observation year, observed value, and data status flags. The sex dimension includes total, male, female, and other categories. The dataset provides standardized, ML-ready data on labor underutilization in Asia, suitable for tabular classification, regression, and time-series forecasting tasks. Data is harmonized using ILO statistical methods for international comparability and is released under the CC-BY-4.0 license.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-eip-dwap-sex-edu-rt-inactivity-rate-by-sex-and-education 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT统计数据库,聚焦于亚洲地区劳动力未充分利用的衡量指标。研究团队通过ILOSTAT的REST API直接提取指标代码为EIP_DWAP_SEX_EDU_RT的原始数据,并依据亚洲国家ISO3代码进行地理过滤。ILOSTAT本身基于国际劳工统计学家会议(ICLS)定义,对各国劳动力调查、家庭收入调查、企业调查及行政记录等微观数据进行标准化处理,以保障跨国可比性。最终数据集收录了39个亚洲国家自1970年至2025年间共计17,694条观测记录,每个观测均标注来源信息,确保可追溯性。
特点
本数据集的核心特色在于其精细的维度划分与丰富的字段设计。除提供不活动率这一核心指标外,数据按性别(总、男、女、其他)和受教育程度进行了细分,支持多维度交叉分析。每个观测值包含国家代码、来源信息、指标标签、时间戳及观测值等20个字段,并附有观测状态标识(如临时、中断等)与各类注释说明,便于用户评估数据质量。数据集的时间跨度长达55年,覆盖东亚、东南亚、南亚、西亚等地区的主要经济体,为区域劳动力市场研究提供了宝贵的纵向对比资源。
使用方法
用户可通过HuggingFace Datasets库便捷调用该数据,使用load_dataset函数即可加载为Pandas DataFrame格式。研究者能按国家代码筛选特定国家的时间序列,或按性别与教育程度分组进行对比分析。支持将数据透视转换为国家-年份矩阵,便于进行面板数据建模。在机器学习应用方面,该数据集适用于表格分类、回归分析及时间序列预测等任务,可构建预测模型分析不活动率的变化趋势。数据采用CC-BY-4.0许可协议,使用时需注明ILO原始来源及Electric Sheep Asia的重组贡献。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)于2025年通过其ILOSTAT数据库整理并发布,后经Electric Sheep Asia重新封装。数据集聚焦于亚洲39个国家1970年至2025年间按性别和教育程度划分的不活动率(%),共计17,694条观测记录。作为劳动利用不足的重要衡量指标,不活动率揭示了劳动力市场中潜在劳动力的闲置状况,对于理解区域经济结构、性别平等及教育投资回报具有深远意义。该数据集通过整合各国劳动力调查、家庭收入调查等多元来源,为亚洲劳动经济学、性别研究和教育政策分析提供了标准化、跨时空的比较基础,尤其推动了对发展中国家劳动力市场动态的系统性探索。
当前挑战
该数据集所解决的领域核心挑战在于量化亚洲地区因性别和教育差异导致的劳动力闲置现象,突破传统失业率指标的局限性,更全面反映劳动利用不足的复杂性。在构建过程中,挑战主要来自三个方面:一是跨国数据源的高度异质性,包括各国调查方法、分类标准不一致,需依赖ILO通过ICLS定义进行协调;二是时间序列中存在中断标记(break in series)和方法论修订(methodology revised),可能导致不同时期数据的可比性下降;三是教育分类的非标准化问题(如nonstandard education level),需借助注释字段追溯原始含义,增加了数据清洗与解析的难度。
常用场景
经典使用场景
该数据集在劳动经济学与区域发展研究中扮演着重要角色,其最典型的应用场景在于对亚洲39个国家1970年至2025年间按性别与教育程度划分的劳动力闲置率进行长时序、跨国别的系统分析。研究者可借助该数据构建面板数据模型,深入探究教育水平提升与女性劳动参与率之间的动态关联,或评估不同经济发展阶段下劳动力市场结构性失业的变化趋势。数据集按性别(男性、女性、总体)和教育层次(如总教育水平)提供的细粒度分类,使其成为检验人力资本理论、性别平等政策效果及劳动力市场弹性等经典议题的理想素材。
解决学术问题
该数据集有效回应了劳动经济学中若干长期存在的学术难题,尤其是关于教育与性别差异对劳动力市场表现的非对称影响。通过提供统一口径下亚洲各国按性别与教育程度分层的经济活动人口比率数据,研究者能够克服以往因各国统计标准不一致而导致比较分析偏误的问题。数据集的存在使得对‘教育能否显著降低女性非活跃率’这一核心命题的因果推断成为可能,同时为评估国际劳工组织倡导的体面劳动目标在亚洲地区的实现进展提供了扎实的量化基础,推动了跨国劳动力市场比较研究的规范化与深度化。
衍生相关工作
该数据集的出现催生了一系列具有影响力的后续研究与实践。基于其提供的标准化长序列数据,研究者已构建起亚洲劳动力闲置率的预测模型,结合宏观经济变量(如GDP增长率、产业结构占比)探索非活跃率的先导性指标;亦有学者将其与ILOSTAT其他指标(如失业率、非正规就业比例)关联分析,形成对劳动力市场多重闲置形态的综合认知框架。此外,Electric Sheep Asia团队对该数据的重新打包与标准化处理,推动了跨数据集联用的可行性,激励了更多以亚洲为聚焦点的劳动统计机器学习项目,例如利用时间序列聚类算法识别具有相似劳动力市场演变模式的国家群组。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务