遇见数据集

electricsheepasia/asia-ilo-luu-xlu4-sex-nb-composite-measure-of-labour-underutilization-lu4-b

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - n<1K tags: - tabular - asia - ilostat - other-measures-of-labour-underutilization - ilo - labour - employment pretty_name: "Composite measure of labour underutilization (LU4) by sex (thousands) | Asia (ILOSTAT)" --- # Composite measure of labour underutilization (LU4) by sex (thousands) | Asia (ILOSTAT) 🌏 **582 observations** · **31 Asia countries** · **1999–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-582-blue) ![countries](https://img.shields.io/badge/countries-31-green) ![years](https://img.shields.io/badge/years-1999–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **582 observations** of `Other measures of labour underutilization` data across **31 Asia countries**, spanning **1999–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=LUU_XLU4_SEX_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Other measures of labour underutilization ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=LUU_XLU4_SEX_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 31 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `CYP` | 78 | 1999 | 2024 | | `VNM` | 51 | 2007 | 2024 | | `LKA` | 42 | 2010 | 2024 | | `KGZ` | 39 | 2011 | 2023 | | `THA` | 39 | 2010 | 2024 | | `TUR` | 30 | 2004 | 2013 | | `PSE` | 30 | 2015 | 2025 | | `BRN` | 27 | 2014 | 2024 | | `JOR` | 24 | 2017 | 2024 | | `PHL` | 21 | 2017 | 2023 | | `IDN` | 21 | 2016 | 2023 | | `GEO` | 18 | 2019 | 2024 | | `MNG` | 18 | 2019 | 2024 | | `KHM` | 15 | 2003 | 2019 | | `AFG` | 15 | 2012 | 2021 | | ... | _16 more countries_ | | | ## Indicators (sample) - `LUU_XLU4_SEX_NB` — Composite measure of labour underutilization (LU4) by sex (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `LUU_XLU4_SEX_NB` | | `indicator.label` | `string` | Indicator name in English | `Composite measure of labour underutil…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `1688.114` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `B` | | `obs_status.label` | `string` | — | `Break in series` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-luu-xlu4-sex-nb-composite-measure-of-labour-underutilization-lu4-b") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "LUU_XLU4_SEX_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="LUU_XLU4_SEX_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "LUU_XLU4_SEX_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_luu_xlu4_sex_nb_composite_measure_of_labour_underutilization_lu4_b_2025, title = {Composite measure of labour underutilization (LU4) by sex (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=LUU_XLU4_SEX_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-luu-xlu4-sex-nb-composite-measure-of-labour-underutilization-lu4-b}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=LUU_XLU4_SEX_NB_

This dataset contains composite measure of labour underutilization (LU4) by sex (thousands) for 31 Asia countries from 1999 to 2025, sourced from the International Labour Organization (ILO) ILOSTAT database. It includes 582 observations covering one distinct indicator (LUU_XLU4_SEX_NB), with details such as country codes, country names, data sources, indicator codes, sex disaggregation (total, male, female), observation years, observed values, and status flags. The data is harmonized using International Conference of Labour Statisticians (ICLS) definitions and is suitable for tabular classification, regression, and time-series forecasting tasks, with annual frequency and best-source selection for data quality.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-luu-xlu4-sex-nb-composite-measure-of-labour-underutilization-lu4-b 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT统计数据库,经由Electric Sheep Asia重新封装而成。数据通过ILOSTAT REST API接口直接获取,依据ILO的ICLS(国际劳工统计学家会议)定义对原始调查微观数据进行协调统一。研究者聚焦于亚洲区域,依据ISO 3166-1 alpha-3国家代码进行筛选,收集了1999年至2025年间31个亚洲国家关于劳动力利用不足综合指标(LU4)的年度数据,共计582条观测记录,并按性别维度进行细分。
特点
该数据集的核心特色在于其聚焦于劳动力利用不足的综合度量(LU4),能够全面反映劳动市场的闲散状况。数据覆盖亚洲31国,跨越27年,时间跨度长且地域范围广。指标采用ILO统一口径,保证了跨国的可比性,同时提供了详细的分类维度,包括性别(男、女、总计)以及丰富的元数据字段,如数据来源、观测状态和断点说明,便于用户深入分析劳动力结构的动态变迁。
使用方法
用户可通过HuggingFace的datasets库快速加载数据,示例代码简洁直观。数据集以表格形式组织,支持多种分析操作:既可针对单一国家(如印度尼西亚)进行筛选,也可对特定指标进行时间序列的可视化展示,更能通过透视表功能灵活构建国家与年份的交叉矩阵,以满足从微观国家维度的比较研究到宏观区域趋势的宏观分析需求,极大地方便了劳动经济学研究者与政策制定者的数据应用。
背景与挑战
背景概述
在劳动经济学与可持续发展目标监测领域,劳动力利用不足的精确度量长期面临概念与方法论挑战。为响应国际劳工组织(ILO)对‘体面工作’指标体系的完善需求,本数据集于2025年由Electric Sheep Asia团队基于ILOSTAT权威数据库二次整理发布,聚焦亚洲31国1999至2025年间复合劳动力利用不足指标(LU4)的性别细分数据,共含582条观测记录。该数据集的核心研究问题在于,通过标准化ICLS定义整合各国劳动力调查微观数据,呈现性别维度的劳动力闲置全貌,为区域劳动力市场结构分析、政策评估及跨国比较提供了珍贵的时间序列基础。其发布填补了公开可获取的亚洲高分辨率劳动力利用不足数据的空白,对发展经济学、性别研究及劳动政策制定具有重要支撑作用,并推动了ILO统计体系与机器学习就绪数据格式的融合。
当前挑战
该领域面临的首要挑战在于度量本身:LU4作为综合指标涵盖失业、时间相关就业不足、潜在劳动力及就业不充分等多重维度,其界定与采集高度依赖各国劳动力调查的规范性,而亚洲各国调查频率、样本框架及统计口径的差异显著影响数据可比性,缺失值与序列断裂(如观测状态标记为B)常见。构建过程中,数据整合须克服多源数据融合、历史序列不连续(如塞浦路斯覆盖1999至2024年,阿富汗仅2012至2021年)及性别细分稀疏等难题;同时,从ILOSTAT REST API提取时需处理指标定义变动、最佳来源选择逻辑,并确保元数据(如源标签、断点注释)的完整传递,以保障数据追溯性与学术严谨性,对数据管线的自动化与质量控制提出了严苛要求。
常用场景
经典使用场景
作为国际劳工组织(ILO)发布的亚洲劳动力利用不足综合指标(LU4)数据集,其经典使用场景聚焦于劳动经济学领域的跨国比较研究。该数据集涵盖31个亚洲国家1999年至2025年的年度观测值,按性别(总、男、女)细致划分,为分析劳动力市场结构性失衡、测度潜在劳动力供给压力提供了扎实的数据基础。研究者可借助该面板数据,运用时间序列分析或面板计量方法,考察亚洲各国劳动力利用不足的长期演变趋势,并评估不同性别群体在劳动市场中的差异化表现,从而深化对劳动力资源错配问题的认知。
实际应用
在实际应用中,该数据集为政府机构、国际组织及智库的劳动力政策制定提供了量化依据。韩国统计厅、亚洲开发银行等机构可据此监测各国劳动力市场健康度,识别劳动力利用不足的高发群体(如女性或特定年龄段),进而设计针对性的就业促进计划与职业培训项目。企业人力资源部门亦可参考此数据评估区域劳动力储备,优化投资选址。此外,该数据集的时间跨度和国别覆盖使其成为联合国可持续发展目标(SDG)第8项(体面工作)进展评估的辅助工具,为全球劳动治理中的区域协调提供了数据参照。
衍生相关工作
该数据集的衍生价值体现在多个经典研究工作上。基于此数据,学者们构建了亚洲劳动力市场脆弱性指数,用以量化经济危机对就业结构的冲击;亦有研究将其与经济增长、教育水平等宏观变量联立,揭示劳动力利用不足与人力资本积累之间的交互效应。ILO自身也常引用此类数据发布年度《世界就业与社会展望》报告,深入剖析区域劳动市场动态。此外,该数据集还催生了基于机器学习的劳动力短缺预测模型,利用时间序列特征预判未来劳动力供给趋势,为前瞻性政策提供工具支撑,凸显了数据基础设施对学术创新与决策科学的双向赋能。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务