遇见数据集

electricsheepasia/asia-ilo-eip-xplf-sex-edu-nb-potential-labour-force-by-sex-and-education-thousa

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 1K<n<10K tags: - tabular - asia - ilostat - other-measures-of-labour-underutilization - ilo - labour - employment pretty_name: "Potential labour force by sex and education (thousands) | Asia (ILOSTAT)" --- # Potential labour force by sex and education (thousands) | Asia (ILOSTAT) 🌏 **7,445 observations** · **32 Asia countries** · **1999–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-7,445-blue) ![countries](https://img.shields.io/badge/countries-32-green) ![years](https://img.shields.io/badge/years-1999–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **7,445 observations** of `Other measures of labour underutilization` data across **32 Asia countries**, spanning **1999–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_XPLF_SEX_EDU_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Other measures of labour underutilization ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EIP_XPLF_SEX_EDU_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 32 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `CYP` | 841 | 1999 | 2024 | | `VNM` | 600 | 2010 | 2024 | | `THA` | 529 | 2010 | 2024 | | `PSE` | 518 | 2012 | 2025 | | `KOR` | 499 | 2000 | 2019 | | `LKA` | 382 | 2010 | 2024 | | `BRN` | 340 | 2014 | 2024 | | `JOR` | 321 | 2017 | 2024 | | `ARM` | 312 | 2007 | 2018 | | `IDN` | 282 | 2015 | 2023 | | `BGD` | 219 | 2013 | 2024 | | `MNG` | 217 | 2019 | 2024 | | `GEO` | 214 | 2019 | 2024 | | `ARE` | 211 | 2017 | 2023 | | `TUR` | 210 | 2000 | 2013 | | ... | _17 more countries_ | | | ## Indicators (sample) - `EIP_XPLF_SEX_EDU_NB` — Potential labour force by sex and education (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EIP_XPLF_SEX_EDU_NB` | | `indicator.label` | `string` | Indicator name in English | `Potential labour force by sex and edu…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `637.031` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:2620` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-eip-xplf-sex-edu-nb-potential-labour-force-by-sex-and-education-thousa") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EIP_XPLF_SEX_EDU_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EIP_XPLF_SEX_EDU_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EIP_XPLF_SEX_EDU_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_eip_xplf_sex_edu_nb_potential_labour_force_by_sex_and_education_thousa_2025, title = {Potential labour force by sex and education (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_XPLF_SEX_EDU_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-eip-xplf-sex-edu-nb-potential-labour-force-by-sex-and-education-thousa}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_XPLF_SEX_EDU_NB_

This dataset contains 7,445 observations of Other measures of labour underutilization data across 32 Asia countries, spanning from 1999 to 2025, covering 1 distinct indicator: Potential labour force by sex and education (thousands). The data is sourced from the International Labour Organizations (ILO) ILOSTAT database, retrieved via REST API and filtered to Asian countries. It includes labour force statistics disaggregated by sex (total, male, female) and education level, designed for analyzing and forecasting labour market conditions in Asia. Organized in tabular format with fields such as country code, year, observed value, source, and quality flags, it is suitable for machine learning tasks like tabular classification, regression, and time-series forecasting.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-eip-xplf-sex-edu-nb-potential-labour-force-by-sex-and-education-thousa 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT中央统计数据库,聚焦于亚洲地区劳动参与不足的测量指标。构建过程中,通过ILOSTAT REST API直接拉取原始数据,并依据国际劳工统计学家会议(ICLS)定义对调查微观数据进行统一协调。数据筛选至32个亚洲国家的ISO3国别代码,覆盖1999至2025年间年度观测值,最终形成包含7,445条记录的表格数据集。每条记录详细记录了性别与教育程度分层的潜在劳动力规模,数据来源经由source.label字段标注,确保了溯源的可追溯性。
特点
本数据集的核心特征在于其多维度分层与区域聚焦的融合。数据以潜在劳动力为观测对象,按性别(总、男、女)及教育程度(整体层次)进行精细划分,揭示了劳动力市场中隐性的未充分利用群体。地理范围涵盖亚洲32国,时间跨度长达27年,为面板数据分析和跨国比较提供了丰富基础。数据集还包含了观测状态标志、来源注释及序列中断说明等元信息,有助于研究者评估数据质量。整体而言,该数据集兼具政策导向性与学术研究价值。
使用方法
数据集以HuggingFace Datasets库提供,可通过load_dataset函数一键加载至Python环境中,并直接转换为pandas DataFrame进行后续操作。用户可根据国别代码(ref_area)筛选特定国家的子集,或通过indicator列锁定单一指标进行时间序列分析。借助pivot_table方法,还可将数据重构为国家×年份的矩阵视图,便于进行面板回归、趋势预测等计量分析。数据采用CC-BY-4.0许可协议,使用时需同时引用ILO原始来源及Electric Sheep Asia的重封装版本。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)统计部门于2025年基于其核心数据库ILOSTAT构建,并由Electric Sheep Asia团队重新封装发布,专注于1999至2025年间亚洲32个国家的潜在劳动力规模(按性别与教育程度细分)。潜在劳动力是衡量劳动力利用不足的重要维度,涵盖了因各种原因未参与就业市场但仍具备劳动能力的群体,对于理解亚洲地区复杂多变的劳动结构、性别差异及教育赋能具有关键意义。作为ILO劳工统计体系中的核心指标之一,该数据集为研究亚洲发展中国家的非正规就业、隐性失业及政策干预效果提供了跨国的、标准化的数据基础,推动了区域劳动经济学与可持续发展目标(SDG)监测的实证研究。
当前挑战
该数据集面临的核心挑战在于如何精准定义与测量“潜在劳动力”这一模糊概念。由于各国采用不同的国家劳动力调查标准(如就业定义、统计口径和调查方法),数据在跨国比较时存在一致性难题,ILO虽通过ICLS定义进行协调,但源数据仍受到调查周期、抽样框架及数据质量的影响,导致部分观测值被标记为“不可靠”。此外,教育分类的差异(如标准与非标准教育水平)增加了维度对齐的复杂性。构建过程中,从多源异构调查数据(如劳动力调查、行政记录)中提取并整合时间序列,处理缺失值、断点和方法修订带来的序列不连续,亦构成技术层面显著挑战。
常用场景
经典使用场景
在劳动经济学与区域发展研究领域,该数据集最为经典的应用场景是对亚洲各国潜在劳动力规模进行跨时空的量化监测与比较分析。凭借跨越1999年至2025年、覆盖32个亚洲国家及性教育与学历分层的7,445条观测记录,研究者可基于时间序列模型追踪潜在劳动力的演变轨迹,或利用面板数据结构剖析不同性别与教育背景群体在劳动力市场参与中的结构性差异。这一数据基础设施为揭示亚洲地区未充分就业的隐性特征提供了宝贵的量化支撑。
解决学术问题
该数据集有效回应了劳动经济学中关于就业不足测度与性别教育分层的学术关切。通过ILOSTAT标准化定义的潜在劳动力指标,研究者得以超越传统的失业率框架,捕捉边缘化劳动群体的潜在供给特征,从而更精准地刻画亚洲劳动力市场的真实弹性与脆弱性。其跨时序、多国别的数据结构为探究经济发展阶段、教育扩张与劳动力参与率之间的交互机制提供了实证依据,推动了关于劳动资源未被充分利用及其社会成本的理论深化。
衍生相关工作
围绕该数据集衍生了一系列具有代表性的学术工作,涵盖了劳动力动态模拟、性别就业差距的计量分析以及教育与劳动参与率关系的因果推断等方向。基于其面板数据结构,研究者构建了包含国家固定效应与时间趋势的回归模型,以揭示不同教育层次人群在潜在劳动力转化率上的异质性。此外,利用时间序列分解技术,相关研究成功分离出亚洲各国潜在劳动力变动的周期性与结构性因素,为该领域后续的理论建构与实证检验奠定了坚实的方法论基础。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务