遇见数据集

electricsheepasia/asia-ilo-eip-xplf-sex-nb-potential-labour-force-by-sex-thousands

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - n<1K tags: - tabular - asia - ilostat - other-measures-of-labour-underutilization - ilo - labour - employment pretty_name: "Potential labour force by sex (thousands) | Asia (ILOSTAT)" --- # Potential labour force by sex (thousands) | Asia (ILOSTAT) 🌏 **798 observations** · **36 Asia countries** · **1999–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-798-blue) ![countries](https://img.shields.io/badge/countries-36-green) ![years](https://img.shields.io/badge/years-1999–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **798 observations** of `Other measures of labour underutilization` data across **36 Asia countries**, spanning **1999–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_XPLF_SEX_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Other measures of labour underutilization ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EIP_XPLF_SEX_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 36 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `CYP` | 78 | 1999 | 2024 | | `PHL` | 57 | 2003 | 2023 | | `VNM` | 51 | 2007 | 2024 | | `KOR` | 45 | 2000 | 2019 | | `LKA` | 42 | 2010 | 2024 | | `TUR` | 42 | 2000 | 2013 | | `PSE` | 39 | 2012 | 2025 | | `KGZ` | 39 | 2011 | 2023 | | `THA` | 39 | 2010 | 2024 | | `ARM` | 36 | 2007 | 2018 | | `BRN` | 27 | 2014 | 2024 | | `ARE` | 24 | 2017 | 2024 | | `IDN` | 24 | 2015 | 2023 | | `JOR` | 24 | 2017 | 2024 | | `MYS` | 18 | 2017 | 2022 | | ... | _21 more countries_ | | | ## Indicators (sample) - `EIP_XPLF_SEX_NB` — Potential labour force by sex (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EIP_XPLF_SEX_NB` | | `indicator.label` | `string` | Indicator name in English | `Potential labour force by sex (thousa…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `637.031` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `B` | | `obs_status.label` | `string` | — | `Break in series` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-eip-xplf-sex-nb-potential-labour-force-by-sex-thousands") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EIP_XPLF_SEX_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EIP_XPLF_SEX_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EIP_XPLF_SEX_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_eip_xplf_sex_nb_potential_labour_force_by_sex_thousands_2025, title = {Potential labour force by sex (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_XPLF_SEX_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-eip-xplf-sex-nb-potential-labour-force-by-sex-thousands}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_XPLF_SEX_NB_

This dataset is titled Potential labour force by sex (thousands) | Asia (ILOSTAT) and is a tabular dataset focusing on labour underutilization in Asia, suitable for tabular classification, regression, and time-series forecasting tasks. It contains 798 observations across 36 Asian countries, spanning the years 1999 to 2025, with 1 distinct indicator: EIP_XPLF_SEX_NB (potential labour force by sex, in thousands). The data is sourced from the International Labour Organization (ILO) ILOSTAT database, retrieved via API and filtered to Asian countries, and is licensed under CC-BY-4.0. The dataset includes columns such as country code (ref_area), country name (ref_area.label), data source (source), indicator code (indicator), sex disaggregation (sex, including total SEX_T, male SEX_M, female SEX_F), year (time), observed value (obs_value), and more, enabling analysis by gender, country, and time. The data is annual frequency and comes with quality caveats (e.g., breaks in series, source selection). It is designed to provide machine learning-ready labour data for Asia, facilitating quick loading and usage via Hugging Faces `load_dataset` function.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-eip-xplf-sex-nb-potential-labour-force-by-sex-thousands 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT统计数据库,通过调用其REST API接口,获取了指标EIP_XPLF_SEX_NB(按性别划分的潜在劳动力,单位:千人)的原始数据,并依据亚洲国家ISO3代码进行地理范围筛选,最终整合为涵盖36个亚洲国家、时间跨度为1999年至2025年的798条观测记录。数据构建过程中,ILOSTAT依据国际劳动统计学家会议(ICLS)定义对原始调查微观数据进行标准化处理,并在source.label字段中标注数据来源,以确保可追溯性。
特点
该数据集的特点在于聚焦亚洲地区劳动力利用不足的测度,提供了单一指标但具备性别维度的精细分解,包含总人口、男性和女性三个类别,便于进行性别差异分析。数据以年度频率发布,覆盖范围广泛,从塞浦路斯到菲律宾等36个国家,时间跨度长达27年。此外,数据集还包含了观测状态标志(如临时性或数据中断)及相关注释信息,为数据质量评估提供了丰富线索,整体呈现结构清晰、来源权威且便于机器读取的优势。
使用方法
研究者可通过HuggingFace的datasets库便捷加载该数据集,仅需一行代码`load_dataset('electricsheepasia/asia-ilo-eip-xplf-sex-nb-potential-labour-force-by-sex-thousands')`即可获取为Parquet格式的数据,并进一步转换为Pandas DataFrame进行深入分析。典型应用包括筛选特定国家(如印度尼西亚IDN)的时序数据、按指标排序后绘制趋势图,或利用透视表功能构建以年份为行、国家为列的交叉矩阵,以支持跨国的面板数据分析与时间序列预测任务。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)于2025年发布,经Electric Sheep Asia团队重新整理并托管于HuggingFace平台,聚焦于亚洲36个国家1999至2025年间潜在劳动力(按性别划分,单位:千人)的统计指标。潜在劳动力作为劳动力利用不足的重要衡量维度之一,反映了那些虽未积极求职但仍有工作意愿的人群规模,对于深入理解区域劳动市场的结构性特征具有关键意义。数据集基于ILOSTAT官方API提取,整合了各国劳动调查数据,并遵循国际劳工统计学家会议(ICLS)定义进行标准化处理,为研究亚洲劳动力市场波动、性别差异及劳动政策效果评估提供了跨时空的可比性基础,在经济社会发展研究领域具备重要价值。
当前挑战
该数据集面临多重挑战。首先,在领域问题层面,潜在劳动力指标在不同国家间的统计口径、调查方法及数据质量存在显著差异,导致跨国比较时需谨慎处理系统性偏差。其次,构建过程中,数据来源涉及多国劳动调查、行政记录及家庭收入调查,整合时需应对时间跨度长(1999-2025)、数据缺失性强及标注不一致的问题,例如部分观测值存在中断标记或方法论修订注释。此外,数据按年度发布,缺乏月度或季度高频信息,限制了短期波动分析的可能性。最后,性别维度的分离虽提供了关键视角,但部分国家数据稀疏,可能影响特定子群体的推断稳健性。
常用场景
经典使用场景
该数据集收录了1999年至2025年间亚洲36个国家的潜在劳动力人口按性别分组的年度观测数据,共计798条记录。在经典使用场景中,研究者常利用此数据集构建跨国的劳动力利用不足指标面板数据,结合时间序列分析方法,揭示亚洲地区潜在劳动力规模的动态演变规律。通过按性别维度进行分解,该数据能够支撑劳动经济学领域关于性别差异在劳动力边缘化群体中表现形态的实证分析,特别适用于评估因经济环境波动或政策干预导致的‘潜在劳动力’——即那些虽未积极求职但有就业意愿的人群——的规模变化趋势。
衍生相关工作
该数据集衍生了一系列以劳动力利用不足为核心议题的经典研究工作。基于ILOSTAT的规范统计框架,学者们利用此类数据开展过亚洲国家潜在劳动力规模与GDP增长率的协整分析、性别差异在劳动力边缘群体中的形成机制研究,以及经济危机前后潜在劳动力数量的脉冲响应建模等。此外,该数据集常与其他劳动力指标如失业率、就业人口比、劳动参与率等联合使用,构建多维劳动力市场健康状况的指数化评估体系。它也为机器学习方法在劳动力预测领域的应用提供了标准化训练素材,催生了基于时序模型与面板回归的劳动力供给趋势预测文献。
数据集最近研究
最新研究方向
该数据集聚焦于亚洲地区潜在劳动力规模的性别差异研究,为劳动经济学中的劳动力利用不充分问题提供了关键数据支持。基于ILOSTAT官方统计,覆盖36个亚洲国家长达26年的时序观测,该数据集在弹性就业、非正规经济及隐性失业等前沿领域具有重要价值。结合近年来全球疫情后劳动力市场复苏、亚洲新兴经济体性别就业差距扩大等热点事件,研究者可利用该数据进行跨性别劳动参与率的动态建模与政策评估,为推动体面劳动目标(SDG 8)及区域劳动力市场均衡发展提供量化依据。其结构化、机器可读的格式亦便于与宏观经济指标联合分析,助力揭示性别维度下的劳动力储备机制与结构性转型路径。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务