遇见数据集

electricsheepasia/asia-ilo-eip-dwap-sex-edu-cbr-rt-inactivity-rate-by-sex-education-and-place-of-birt

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 1K<n<10K tags: - tabular - asia - ilostat - international-migrant-stock - ilo - labour - employment pretty_name: "Inactivity rate by sex, education and place of birth (%) | Asia (ILOSTAT)" --- # Inactivity rate by sex, education and place of birth (%) | Asia (ILOSTAT) 🌏 **5,278 observations** · **19 Asia countries** · **1999–2024** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-5,278-blue) ![countries](https://img.shields.io/badge/countries-19-green) ![years](https://img.shields.io/badge/years-1999–2024-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **5,278 observations** of `International migrant stock` data across **19 Asia countries**, spanning **1999–2024**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_DWAP_SEX_EDU_CBR_RT) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** International migrant stock ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EIP_DWAP_SEX_EDU_CBR_RT` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 19 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `CYP` | 1,180 | 1999 | 2024 | | `TUR` | 774 | 2000 | 2024 | | `ISR` | 701 | 2012 | 2024 | | `ARM` | 426 | 2001 | 2023 | | `BRN` | 420 | 2014 | 2024 | | `TLS` | 268 | 2001 | 2022 | | `KHM` | 244 | 2009 | 2021 | | `IDN` | 180 | 2017 | 2022 | | `MNG` | 177 | 2019 | 2024 | | `MDV` | 168 | 2014 | 2019 | | `IRQ` | 149 | 2007 | 2021 | | `MMR` | 110 | 2014 | 2015 | | `ARE` | 98 | 2022 | 2023 | | `LAO` | 85 | 2017 | 2022 | | `TJK` | 80 | 2007 | 2009 | | ... | _4 more countries_ | | | ## Indicators (sample) - `EIP_DWAP_SEX_EDU_CBR_RT` — Inactivity rate by sex, education and place of birth (%) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:6361` | | `source.label` | `string` | Source name in English | `HIES - Households Living Conditions S…` | | `indicator` | `string` | ILOSTAT indicator code | `EIP_DWAP_SEX_EDU_CBR_RT` | | `indicator.label` | `string` | Indicator name in English | `Inactivity rate by sex, education and…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `classif2` | `string` | Second classification variable where applicable | `CBR_BIR_TOTAL` | | `classif2.label` | `string` | — | `Place of birth: Total` | | `time` | `int64` | Observation year | `2014` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `48.275` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:2620` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-eip-dwap-sex-edu-cbr-rt-inactivity-rate-by-sex-education-and-place-of-birt") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EIP_DWAP_SEX_EDU_CBR_RT"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EIP_DWAP_SEX_EDU_CBR_RT") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EIP_DWAP_SEX_EDU_CBR_RT"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_eip_dwap_sex_edu_cbr_rt_inactivity_rate_by_sex_education_and_place_of_birt_2024, title = {Inactivity rate by sex, education and place of birth (%) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2024}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_DWAP_SEX_EDU_CBR_RT}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-eip-dwap-sex-edu-cbr-rt-inactivity-rate-by-sex-education-and-place-of-birt}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_DWAP_SEX_EDU_CBR_RT_

This dataset contains 5,278 observations of inactivity rate statistics disaggregated by sex, education, and place of birth across 19 Asia countries from 1999 to 2024. The core indicator is EIP_DWAP_SEX_EDU_CBR_RT (Inactivity rate by sex, education and place of birth %). Data is sourced from the International Labour Organization (ILO) ILOSTAT database, retrieved via its REST API and filtered for Asia ISO3 country codes. The dataset includes detailed disaggregation dimensions such as sex (total, male, female), education (aggregate levels), and place of birth (total), along with country codes, year, observed values, data sources, and quality flags. Data is provided at annual frequency and is suitable for tabular classification, regression, and time-series forecasting tasks.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-eip-dwap-sex-edu-cbr-rt-inactivity-rate-by-sex-education-and-place-of-birt 数据集图片
构建方式
该数据集源于国际劳工组织(ILO)旗下的ILOSTAT统计数据库,聚焦于亚洲19个国家中,按性别、教育程度及出生地划分的劳动参与不活跃率(%)。数据通过ILOSTAT提供的REST API直接拉取,并依据亚洲地区的ISO3国家代码进行过滤,最终得到5,278条观测记录,时间跨度从1999年至2024年。ILO在构建原始数据时,依据国际劳工统计学家会议(ICLS)的定义对各国劳动力调查的微观数据进行统一协调与标准化,确保了跨国数据的可比性。每条观测记录均包含对数据来源的标注,以增强可追溯性,而Electric Sheep Asia团队则对原始数据进行了整合与重装,使其更便于机器学习使用。
特点
数据集的核心特色在于其精细的维度划分,不仅涵盖性别(男、女、总计),还通过两个分类变量(如教育水平与出生地)对不活跃率进行深层剖析,揭示了劳动参与情况的复杂社会结构。其地理覆盖范围广泛,囊括了塞浦路斯、土耳其、以色列、印度尼西亚等19个亚洲经济体,数据来源的多样性(如家庭收入支出调查、劳动力调查等)均被清晰记录。此外,数据集附带了观测状态标志和注释信息,用于标识数据可靠性、方法修订及非标准分类等潜在问题,为用户提供透明的数据质量评估依据。
使用方法
使用者可通过HuggingFace Datasets库的`load_dataset()`函数轻松加载该数据集,并将其转换为Pandas DataFrame进行后续分析。典型的使用场景包括按国家进行数据切片,如筛选印度尼西亚的观测值;或对单一指标进行时间序列分析,如按年份排序后绘制不活跃率的变化趋势。此外,用户还可利用透视表功能将数据重组为国家-年份的矩阵形式,便于进行面板数据分析或跨国家比较。所有操作均基于标准化列名(如`ref_area`、`time`、`obs_value`),代码简洁高效,适合快速开展劳动经济学或人口学领域的实证研究。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)统计数据库ILOSTAT创建,经Electric Sheep Asia于2024年重新整理发布,聚焦亚洲19个国家1999至2024年间按性别、教育程度和出生地划分的劳动不活跃率。核心研究问题在于揭示亚洲地区人口在劳动力市场之外的分布特征,尤其是教育、性别与迁移背景如何影响个体的经济活动状态。作为ILOSTAT体系中关键的社会劳动力指标之一,该数据集为跨国比较、劳动力政策评估及可持续发展目标(SDGs)中的体面劳动议程提供了标准化、可复用的数据基础。其影响力体现在促进亚洲区域劳动经济学研究的数据透明化,并支持机器学习模型对时间序列与结构化分类任务的建模需求。
当前挑战
数据集面临的领域挑战主要源于劳动不活跃率本身的复杂性,需同时解析性别差异、教育分层与出生地背景的交互对劳动力参与决策的影响,这比单一维度指标更为难以建模。构建过程中,数据的来源多样性与质量参差带来显著困难:ILOSTAT虽致力于协调各国劳动力调查、住户收入调查及行政记录,但不同调查方法的定义差异(如国际劳动统计学家会议定义的应用偏差)以及部分国家数据存在不连续、标记为‘不可靠’或存在方法修订导致的时间序列断裂,使得跨年、跨国比较的效度受限。此外,分类变量(如性别、教育级别)的稀疏性及某些年份观测值不足,也增加了数据分析与模型训练中特征工程的挑战。
常用场景
经典使用场景
该数据集是研究亚洲地区劳动力市场非活跃人口结构的核心资源,收录了1999年至2024年间19个亚洲国家按性别、教育程度和出生地分组的非活动率观测值。经典使用场景集中于构建面板数据模型,以揭示经济发展阶段、社会保障体系与劳动参与率之间的动态关联。研究者常利用该数据集进行跨国比较分析,探讨女性群体在不同教育层级下的就业障碍,或追踪移民人口与本地居民在劳动市场退出概率上的差异趋势。数据的时间跨度与维度拆解使其成为时间序列回归和因果推断的理想基础,尤其适用于分析政策干预如教育改革或就业促进计划对非活跃群体的长期影响。
解决学术问题
该数据集着力破解劳动经济学中关于非活跃劳动力群体的结构性认知困境,传统研究常因数据粗糙而难以区分自愿性退出与被迫性失业。通过提供性别、教育程度与出生地的三重交叉分类,该数据助力学者分离人口特征对劳动参与率的异质性作用,从而澄清教育回报率在不同性别群体中的衰减路径。此外,数据集帮助解决移民融入研究中的关键痛点——量化外来人口因语言、技能或歧视而沦为经济非活跃者的概率,为评估移民政策的社会经济效益奠定了基石。其在学术上的深远意义在于,将亚洲地区碎片化的劳动统计整合至可复现的分析框架内,推动区域比较社会学从描述性叙事迈向定量建模的新高度。
衍生相关工作
围绕该数据集已衍生出一系列推动劳动经济学方法论革新的经典工作。在特征工程层面,研究者将性别-教育-出生地的组合字段编码为交互特征,结合梯度提升或随机森林模型,实现了非活动率预测精度的显著跃升;在迁移学习领域,该数据被用作预训练框架的亚洲区域基准,辅助模型在数据稀疏的西亚与东南亚国家间进行知识迁移。此外,基于该数据集构建的时序神经网络在劳动力供给预警领域产生了广泛影响,其预测结果被整合进多国劳动部的政策仿真平台。值得关注的是,数据集的标准化格式催生了跨数据集集成的研究范式——与ILOSTAT其他就业、薪酬指标融合后,学界得以建立从失业至退出市场的完整马尔可夫状态转移模型,深化了对劳动流动性微观机制的理解。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务