遇见数据集

electricsheepasia/asia-ilo-emp-xtru-sex-geo-mts-nb-time-related-underemployment-by-sex-rural-urban-ar

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 1K<n<10K tags: - tabular - asia - ilostat - time-related-underemployment - ilo - labour - employment pretty_name: "Time-related underemployment by sex, rural / urban area and marital status (thousands) | Asia (ILOSTAT)" --- # Time-related underemployment by sex, rural / urban area and marital status (thousands) | Asia (ILOSTAT) 🌏 **5,407 observations** · **25 Asia countries** · **1996–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-5,407-blue) ![countries](https://img.shields.io/badge/countries-25-green) ![years](https://img.shields.io/badge/years-1996–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **5,407 observations** of `Time-related underemployment` data across **25 Asia countries**, spanning **1996–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_GEO_MTS_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Time-related underemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EMP_XTRU_SEX_GEO_MTS_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 25 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `CYP` | 589 | 1999 | 2020 | | `VNM` | 439 | 2010 | 2024 | | `THA` | 405 | 2010 | 2024 | | `KHM` | 391 | 1996 | 2023 | | `LKA` | 377 | 2010 | 2024 | | `PAK` | 351 | 2006 | 2025 | | `KOR` | 351 | 2013 | 2025 | | `MNG` | 316 | 2013 | 2024 | | `TUR` | 287 | 2004 | 2013 | | `PSE` | 279 | 2015 | 2022 | | `BRN` | 243 | 2014 | 2024 | | `PHL` | 235 | 2017 | 2023 | | `IDN` | 216 | 2016 | 2023 | | `JOR` | 165 | 2017 | 2024 | | `AFG` | 137 | 2014 | 2021 | | ... | _10 more countries_ | | | ## Indicators (sample) - `EMP_XTRU_SEX_GEO_MTS_NB` — Time-related underemployment by sex, rural / urban area and marital status (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EMP_XTRU_SEX_GEO_MTS_NB` | | `indicator.label` | `string` | Indicator name in English | `Time-related underemployment by sex, …` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `GEO_COV_NAT` | | `classif1.label` | `string` | — | `Area type: National` | | `classif2` | `string` | Second classification variable where applicable | `MTS_AGGREGATE_TOTAL` | | `classif2.label` | `string` | — | `Marital status (Aggregate): Total` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `588.672` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-emp-xtru-sex-geo-mts-nb-time-related-underemployment-by-sex-rural-urban-ar") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EMP_XTRU_SEX_GEO_MTS_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EMP_XTRU_SEX_GEO_MTS_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EMP_XTRU_SEX_GEO_MTS_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_emp_xtru_sex_geo_mts_nb_time_related_underemployment_by_sex_rural_urban_ar_2025, title = {Time-related underemployment by sex, rural / urban area and marital status (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_GEO_MTS_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-emp-xtru-sex-geo-mts-nb-time-related-underemployment-by-sex-rural-urban-ar}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_GEO_MTS_NB_

This dataset contains 5,407 observations of time-related underemployment data across 25 Asia countries, spanning 1996–2025, covering 1 distinct indicator (EMP_XTRU_SEX_GEO_MTS_NB). The data is disaggregated by sex (total, male, female), rural/urban area, and marital status, measured in thousands. It is sourced from the International Labour Organization (ILO) ILOSTAT database, retrieved via REST API and filtered to Asia ISO3 country codes, with harmonization based on International Conference of Labour Statisticians (ICLS) definitions. The schema includes fields such as country code, source, indicator, classification variables, observation year, value, and status flags. The dataset is suitable for tabular classification, regression, and time-series forecasting tasks, released under the CC-BY-4.0 license.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-emp-xtru-sex-geo-mts-nb-time-related-underemployment-by-sex-rural-urban-ar 数据集图片
构建方式
该数据集源自国际劳工组织中央统计数据库ILOSTAT,以REST API方式直接采集指标EMP_XTRU_SEX_GEO_MTS_NB的原始记录,并按照亚洲ISO3国家代码进行筛选与截取。原始微观调查数据经国际劳工统计学家会议定义体系协调统一,来源涵盖各国劳动力调查、家庭收入调查及行政记录,每条观测均保留来源标签以便溯源。数据经Electric Sheep Asia重新打包为Parquet格式,在保持原始字段结构与观测状态标记的前提下完成规范化整理,形成覆盖多国多年份的面板型数据资源。
特点
数据集涵盖25个亚洲国家、5407条观测记录,时间跨度为1996年至2025年,以年度频率发布。核心变量涵盖性别、城乡区域与婚姻状况等多重分类维度,并附有来源代码、指标代码、观测状态及系列断裂说明等元数据字段。观测值以千人为计量单位,部分记录标注为不可靠或临时性估计,分类维度字段仅在指标发布对应细分时非空,整体呈现稀疏型面板结构,适合分析亚洲地区与时间相关的就业不足现象及其性别和城乡差异。
使用方法
研究者可借助HuggingFace datasets库以load_dataset函数直接加载数据并转换为Pandas数据框,进而按国家代码、指标代码或年份进行筛选与分组。典型操作包括提取单一国家的时序记录、对特定指标按时间排序绘制趋势图,以及将数据透视为国家与年份的交叉矩阵以开展跨国比较。该数据集亦可用于表格分类、回归及时间序列预测等机器学习任务,使用时须遵循CC-BY-4.0许可协议并同时引用国际劳工组织原始来源与Electric Sheep Asia的再打包工作。
背景与挑战
背景概述
国际劳工组织(ILO)自1919年成立以来,始终致力于全球劳动统计的标准化与推广,其ILOSTAT数据库作为全球劳动统计的权威来源,覆盖200余个经济体,为就业、失业、工资及劳动时间等关键指标提供跨国可比数据。亚洲地区作为全球劳动力最为密集且经济结构快速变迁的区域,其就业不足问题长期缺乏系统性、高颗粒度的时序数据。在此背景下,Electric Sheep Asia基于ILOSTAT官方API,于2025年重新封装并发布了覆盖25个亚洲国家、1996至2025年时间相关就业不足的观测数据集,包含5,407条记录,按性别、城乡及婚姻状况等多维度 disaggregation,为劳动经济学、发展研究及政策评估提供了精细化的数据基础,对探索亚洲劳动力市场动态具有显著的学术与政策价值。
当前挑战
该数据集所应对的领域问题在于准确量化并追踪亚洲地区时间相关就业不足的演变趋势,这一问题因各国劳动统计口径差异、数据收集能力悬殊及非正规经济广泛存在而极具复杂性。在构建过程中,挑战尤为突出:部分国家(如阿富汗、巴基斯坦)数据年份断裂且覆盖率低,样本代表性受限;不同来源的劳动力调查在定义与分类标准上存在异质性,尽管ILO已进行调和,但观测状态标记(如临时、不可靠)仍揭示数据质量参差不齐;性别、城乡及婚姻状况的多维交叉 disaggregation 进一步加剧了样本稀疏问题,导致某些国家—年份—子群体组合存在缺失,为时间序列建模与因果推断带来潜在偏误风险。
常用场景
经典使用场景
在国际劳动统计领域,时间相关就业不足是衡量劳动力市场供需失衡的关键指标。该数据集汇聚了25个亚洲国家1996至2025年间按性别、城乡及婚姻状况分列的时间相关就业不足数据,最经典的使用场景是构建面板数据回归模型,探究性别差异、城乡二元结构以及婚姻状态对就业不足率的交互影响。研究者亦可利用时间序列分析方法,追踪各国就业不足的动态演变轨迹,识别经济周期波动与劳动力市场政策调整之间的关联,为区域劳动市场比较研究提供坚实的数据基础。
实际应用
在实际应用层面,该数据集为国际组织、政府部门及智库机构提供了可靠的决策参考。政策制定者可依据分性别、分城乡的就业不足率变化,评估劳动力市场政策的靶向效果,优化就业服务资源配置。企业人力资源部门可借助国别时间序列数据,研判目标市场的人力供给充裕度与用工成本趋势,辅助海外投资布局。此外,非政府组织亦可运用该数据识别脆弱群体,设计针对农村女性或特定婚姻状态劳动者的技能提升与就业支持项目,推动体面劳动议程的落地。
衍生相关工作
围绕该数据集及其源指标,学界与政策研究机构衍生出一系列经典工作。比较劳动制度研究利用该数据验证了就业保护立法对工时不足的抑制效应;性别经济学文献据此探讨了婚姻状况与女性劳动参与质量之间的关联;区域发展研究则将其与GDP增长、产业结构转型等宏观变量结合,分析亚洲经济体就业不足的结构性诱因。这些工作不仅拓展了国际组织统计数据的学术价值,也为后续开发更细粒度的劳动力市场监测工具奠定了方法论基础。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务