遇见数据集

electricsheepasia/asia-ilo-emp-xtru-sex-age-nb-time-related-underemployment-by-sex-and-age-thousa

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - asia - ilostat - time-related-underemployment - ilo - labour - employment pretty_name: "Time-related underemployment by sex and age (thousands) | Asia (ILOSTAT)" --- # Time-related underemployment by sex and age (thousands) | Asia (ILOSTAT) 🌏 **15,821 observations** · **37 Asia countries** · **1990–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-15,821-blue) ![countries](https://img.shields.io/badge/countries-37-green) ![years](https://img.shields.io/badge/years-1990–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **15,821 observations** of `Time-related underemployment` data across **37 Asia countries**, spanning **1990–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_AGE_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Time-related underemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EMP_XTRU_SEX_AGE_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 37 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `CYP` | 1,194 | 1999 | 2024 | | `TUR` | 996 | 2004 | 2024 | | `IRN` | 960 | 2005 | 2024 | | `AZE` | 858 | 2000 | 2022 | | `VNM` | 816 | 2007 | 2024 | | `THA` | 789 | 1991 | 2024 | | `KHM` | 778 | 1996 | 2023 | | `LKA` | 753 | 2009 | 2024 | | `KOR` | 672 | 2012 | 2025 | | `PHL` | 655 | 1990 | 2023 | | `SGP` | 639 | 2009 | 2024 | | `PAK` | 592 | 2000 | 2025 | | `KGZ` | 589 | 2010 | 2023 | | `MNG` | 584 | 2003 | 2024 | | `ISR` | 467 | 2008 | 2024 | | ... | _22 more countries_ | | | ## Indicators (sample) - `EMP_XTRU_SEX_AGE_NB` — Time-related underemployment by sex and age (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EMP_XTRU_SEX_AGE_NB` | | `indicator.label` | `string` | Indicator name in English | `Time-related underemployment by sex a…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHADULT_YGE15` | | `classif1.label` | `string` | — | `Age (Youth, adults): 15+` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `588.672` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C6:2320` | | `note_classif.label` | `string` | — | `Nonstandard age group: Excluding ages…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-emp-xtru-sex-age-nb-time-related-underemployment-by-sex-and-age-thousa") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EMP_XTRU_SEX_AGE_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EMP_XTRU_SEX_AGE_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EMP_XTRU_SEX_AGE_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_emp_xtru_sex_age_nb_time_related_underemployment_by_sex_and_age_thousa_2025, title = {Time-related underemployment by sex and age (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_AGE_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-emp-xtru-sex-age-nb-time-related-underemployment-by-sex-and-age-thousa}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_AGE_NB_

This dataset, titled Time-related underemployment by sex and age (thousands) | Asia (ILOSTAT), is a statistical dataset focusing on underemployment in Asia. It contains 15,821 observations across 37 Asian countries, spanning the years 1990 to 2025. The core indicator is Time-related underemployment, specifically the number of underemployed persons (in thousands) disaggregated by sex and age, with the indicator code EMP_XTRU_SEX_AGE_NB. The data is sourced from the ILOSTAT database of the International Labour Organization (ILO), a leading global source for labour statistics that harmonizes data from national labour force surveys, household income surveys, and other sources. The dataset has been repackaged to provide a unified, machine-learning-ready format. It is structured in tabular form with columns including country code, country name, data source, indicator, sex (total, male, female), age classification, observation year, observed value, observation status, and more. The data is annual and suitable for tasks such as tabular classification, tabular regression, and time-series forecasting. The dataset is released under the cc-by-4.0 license, and users are required to cite both the original ILO source and the repackaging by Electric Sheep Asia.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-emp-xtru-sex-age-nb-time-related-underemployment-by-sex-and-age-thousa 数据集图片
构建方式
该数据集依托国际劳工组织(ILO)中央统计数据库ILOSTAT的权威数据源,通过REST API接口直接提取指标代码为EMP_XTRU_SEX_AGE_NB的原始记录,并依据ISO3国家代码筛选出亚洲地区37个经济体的相关观测。ILOSTAT采用国际劳工统计学家会议(ICLS)标准定义对各国劳动力调查、家庭收入调查等微观数据进行统一调和,数据集在重打包过程中保留了来源标签、分类注释及指标注释等元数据字段,以保障数据溯源与质量评估的可行性。
特点
数据集涵盖1990年至2025年亚洲37国的15,821条年度观测,以千人为计量单位系统记录了与时间相关的就业不足状况。数据按性别(总计、男性、女性)和年龄组进行分类解聚,包含ref_area、source、indicator、obs_value、obs_status等结构化字段,并附有详尽的分类注释与来源说明。观测状态标识可辅助识别临时性或不稳定估计值,多维度元数据为精细化分析提供了充分的透明度与可靠性保障。
使用方法
研究者可通过HuggingFace的datasets库以load_dataset()函数便捷加载该数据集,并转换为Pandas数据框进行后续处理。典型应用包括按ref_area字段筛选特定国家子集、按indicator与time排序生成单一指标的时间序列趋势图、以及利用pivot_table方法构建国家与年份的交叉矩阵。数据格式规整、字段语义清晰,适用于表格分类、回归分析及时间序列预测等多种机器学习任务场景。
背景与挑战
背景概述
时间相关不充分就业作为衡量劳动力市场健康程度的关键指标,反映劳动者工时不足且寻求额外工作的意愿,其监测数据对于评估体面劳动实现进程至关重要。国际劳工组织(ILO)依托其成员国劳动力调查体系,长期编制全球劳动力统计数据库(ILOSTAT),并依据国际劳工统计学家会议(ICLS)标准对原始微观数据进行协调与标准化。在此背景下,Electric Sheep Asia 于2025年对ILOSTAT中亚洲区域的时间相关不充分就业指标(EMP_XTRU_SEX_AGE_NB)进行重新封装,形成覆盖37个亚洲国家或地区、1990至2025年间共计15,821条观测记录的数据集。该数据集延续了ILO官方统计的权威性,为亚洲劳动力市场性别与年龄维度下的不充分就业研究提供了结构化、可即时加载的机器学习就绪型数据资源。
当前挑战
该数据集所应对的领域问题在于,时间相关不充分就业的跨国与跨期比较长期受制于各国劳动力调查在问卷设计、工时门槛及参考周期上的异质性,导致统计口径难以直接对齐,而亚洲区域内部经济发展阶段与就业形态的显著差异进一步加剧了这一问题。在构建过程中,数据面临若干具体挑战:其一,各国调查年份不连续且覆盖区间参差不齐,部分国家仅有个别年份观测,造成时间序列分析中的缺失值与样本不平衡;其二,性别与年龄分类维度仅在指标发布对应细分时方为非空,导致细粒度分析受到系统性限制;其三,部分观测值附有临时性、不可靠或序列中断等状态标记,需在建模中审慎处理;其四,原始数据经ILO最佳来源筛选后仍可能存在来源切换所引致的水平偏移,对趋势估计构成干扰。
常用场景
经典使用场景
在劳动经济学与人口统计学的实证研究中,该数据集最为经典的用途在于刻画亚洲地区时间相关不充分就业的时序演进与性别—年龄维度差异。研究者常以国家为截面单元、年份为时间轴,构建面板数据模型,考察不充分就业率在经济周期波动、产业结构转型与劳动力市场制度变迁中的动态响应。借助`sex`与`classif1`等分类变量,可精确剥离出青年群体、成年男性与成年女性等子样本的异质性轨迹,进而识别就业质量恶化的高发人群与时点,为跨国比较提供统一口径的量化基础。
实际应用
在政策实践层面,该数据集为国际组织与各国劳工主管部门监测可持续发展目标中体面劳动议程的落实进展提供了关键指标来源。就业服务机构可据此识别不充分就业高发的国别与人群,优化职业培训与工时匹配政策;宏观经济部门可将该序列纳入劳动力市场景气监测体系,辅助预判消费疲软与社会保障支出压力。对于跨国企业与人力资源研究机构而言,该数据亦可支撑区域用工环境评估与劳动力供给质量分析,服务于投资布局与合规决策。
衍生相关工作
围绕该数据集及ILOSTAT同源指标,已衍生出一系列具有代表性的研究工作。劳动经济学文献中常见将其与失业率、劳动参与率及非正规就业指标联立,构建劳动市场松弛度的综合测度框架;计量方法层面,研究者利用其时序特征开展结构性断点检验与贝叶斯层次模型估计,以处理小样本国家的数据稀疏问题。此外,国际劳工组织的年度就业与社会展望报告、亚洲开发银行国别诊断以及若干机器学习驱动的劳动力市场预测竞赛,均以该数据或其上游API作为基准输入,推动了就业质量测度从描述统计向预测建模的延伸。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务