遇见数据集

electricsheepasia/asia-ilo-emp-xtru-sex-edu-nb-time-related-underemployment-by-sex-and-education

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 1K<n<10K tags: - tabular - asia - ilostat - time-related-underemployment - ilo - labour - employment pretty_name: "Time-related underemployment by sex and education (thousands) | Asia (ILOSTAT)" --- # Time-related underemployment by sex and education (thousands) | Asia (ILOSTAT) 🌏 **8,900 observations** · **31 Asia countries** · **1996–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-8,900-blue) ![countries](https://img.shields.io/badge/countries-31-green) ![years](https://img.shields.io/badge/years-1996–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **8,900 observations** of `Time-related underemployment` data across **31 Asia countries**, spanning **1996–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_EDU_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Time-related underemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EMP_XTRU_SEX_EDU_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 31 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `IRN` | 948 | 2005 | 2024 | | `CYP` | 944 | 1999 | 2024 | | `VNM` | 643 | 2010 | 2024 | | `KOR` | 587 | 2012 | 2025 | | `THA` | 563 | 2010 | 2024 | | `KHM` | 517 | 1996 | 2023 | | `PAK` | 484 | 2006 | 2025 | | `LKA` | 482 | 2010 | 2024 | | `MNG` | 482 | 2010 | 2024 | | `TUR` | 423 | 2004 | 2024 | | `PSE` | 332 | 2015 | 2025 | | `BRN` | 330 | 2014 | 2024 | | `IDN` | 283 | 2016 | 2023 | | `BGD` | 215 | 2013 | 2024 | | `GEO` | 191 | 2019 | 2024 | | ... | _16 more countries_ | | | ## Indicators (sample) - `EMP_XTRU_SEX_EDU_NB` — Time-related underemployment by sex and education (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EMP_XTRU_SEX_EDU_NB` | | `indicator.label` | `string` | Indicator name in English | `Time-related underemployment by sex a…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `588.672` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:2620` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-emp-xtru-sex-edu-nb-time-related-underemployment-by-sex-and-education") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EMP_XTRU_SEX_EDU_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EMP_XTRU_SEX_EDU_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EMP_XTRU_SEX_EDU_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_emp_xtru_sex_edu_nb_time_related_underemployment_by_sex_and_education_2025, title = {Time-related underemployment by sex and education (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_EDU_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-emp-xtru-sex-edu-nb-time-related-underemployment-by-sex-and-education}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_EDU_NB_

This dataset contains 8,900 observations of time-related underemployment data across 31 Asia countries, spanning the years 1996 to 2025, covering one distinct indicator: EMP_XTRU_SEX_EDU_NB (Time-related underemployment by sex and education in thousands). The data is sourced from ILOSTAT, the International Labour Organizations central statistics database, which harmonizes raw survey microdata using International Conference of Labour Statisticians definitions. The dataset includes columns such as country codes, country names, source codes, indicator codes, sex disaggregation, education classification, observation year, observed values, status flags, and notes. It has been repackaged by Electric Sheep Asia as part of a unified, ML-ready data layer for Asia on HuggingFace, facilitating easy access and analysis for researchers and developers.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-emp-xtru-sex-edu-nb-time-related-underemployment-by-sex-and-education 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的中央统计数据库ILOSTAT,经由Electric Sheep Asia从ILOSTAT REST API直接抽取指标EMP_XTRU_SEX_EDU_NB的原始记录,并依据ISO3国家代码筛选出亚洲地区数据。原始调查微观数据由各国劳动力调查、家庭收入调查等渠道汇集,ILO统计部门按照国际劳工统计学家会议(ICLS)定义进行标准化调和,确保跨国可比性。数据以年度频率呈现,最终形成覆盖31个亚洲国家、时间跨度为1996至2025年的8900条观测记录,并以Parquet格式封装发布。
使用方法
研究者可通过HuggingFace的datasets库以load_dataset函数加载数据集,并转换为Pandas数据框进行灵活操作。典型用法包括筛选特定国家(如印度尼西亚)、按指标代码提取单一时间序列并绘制折线图,或利用透视表将数据重塑为国家×年份矩阵以开展跨国比较。数据集支持表格分类、回归及时间序列预测等任务,用户需注意观测状态标志与分类注释所提示的数据质量信息,并遵循CC-BY-4.0许可协议,在使用时同时引用ILO原始来源与Electric Sheep Asia的再包装工作。
背景与挑战
背景概述
国际劳工组织长期致力于全球劳动力市场统计体系的构建与完善,其ILOSTAT数据库作为全球劳动统计的权威来源,为政策制定与学术研究提供了坚实的数据基础。在此背景下,Electric Sheep Asia于2025年对ILOSTAT原始数据进行系统性重封装,发布了涵盖亚洲31个国家、时间跨度为1996至2025年、共计8900条观测值的时间相关就业不足数据集。该数据集围绕性别与教育维度展开,核心研究问题在于揭示亚洲地区劳动力未充分就业的结构性特征与演变趋势,为劳动经济学、发展研究及社会政策分析提供了可复用的机器学习就绪型数据资源,对推动区域劳动市场实证研究具有重要价值。
当前挑战
该数据集所应对的领域问题在于时间相关就业不足的精确测度与跨国可比性构建——就业不足本身是一个涉及工时、意愿与收入多维交织的复杂概念,不同国家劳动统计口径的差异使得跨国比较面临显著的方法论障碍。构建过程中的核心挑战则体现为:原始调查数据的质量参差不齐,部分观测值被标记为暂定或不可靠;亚洲各国劳动力调查的覆盖年份与频率高度不均衡,导致面板数据存在大量缺失;教育分类体系与性别维度的非标准化处理增加了数据清洗与整合的复杂度;此外,序列中断与方法论修订等结构性变化进一步削弱了时间序列的连续性与一致性。
常用场景
经典使用场景
在劳动经济学与人口统计学的交叉领域,时间相关就业不足作为衡量劳动力未充分利用程度的关键指标,始终是学术探究与政策研判的焦点。该数据集汇聚了1996至2025年间亚洲31个国家的8900条观测记录,按性别与教育水平细分,为研究者构建跨国别、长时序的面板数据分析提供了经典素材。借助这一数据资源,学者可深入刻画亚洲地区就业不足的性别差异与教育梯度,揭示不同受教育群体在劳动力市场中的边缘化程度,并据此开展时间序列预测与横截面比较研究,从而系统把握亚洲劳动力市场结构的动态演变。
解决学术问题
该数据集有效回应了劳动经济学中关于就业质量与人力资本回报的若干核心学术议题。以往研究常受限于跨国就业不足数据的碎片化与口径不一,难以进行严谨的横向比较。此数据集依托国际劳工组织的标准化框架,统一了年龄、教育及性别分类,使学者得以检验教育扩张是否真正缓解了就业不足,以及性别鸿沟在不同经济发展阶段的收敛或分化趋势。其意义在于为劳动力市场分层理论、人力资本理论及性别经济学提供了可复现的实证基础,推动了亚洲就业不足问题的量化研究向纵深发展。
实际应用
在政策制定与劳动力市场监测的实际场景中,该数据集展现出广泛的实用价值。各国劳工部门与国际组织可借助其按性别与教育水平细分的就业不足数据,精准识别高风险群体,例如低学历女性或青年毕业生,从而设计更具针对性的技能培训与就业促进计划。同时,该数据集可用于构建预警指标,评估经济波动对就业质量的冲击,辅助社会保障政策的动态调整。对于企业人力资源规划而言,理解不同教育层级劳动力的未充分就业状况,亦有助于优化招聘策略与岗位配置,提升整体劳动效率。
数据集最近研究
最新研究方向
在全球劳动力市场深刻变革的背景下,亚洲地区时间相关不充分就业的性别与教育维度差异日益成为劳动经济学与发展研究的前沿议题。围绕该数据集的最新研究,学界正着力于运用面板数据模型与时序预测方法,探究教育层级提升对女性不充分就业的缓释效应,并结合国际劳工组织关于体面劳动与可持续发展目标的政策框架,评估非正规经济扩张、数字化转型及后疫情时代复苏不均衡对就业质量的冲击。相关成果为亚洲各国优化教育资源配置、缩小性别就业差距以及制定精准劳动力市场干预政策提供了关键实证依据。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务