遇见数据集

electricsheepasia/asia-ilo-eip-wdis-sex-geo-mts-nb-discouraged-job-seekers-by-sex-rural-urban-area-an

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 1K<n<10K tags: - tabular - asia - ilostat - other-measures-of-labour-underutilization - ilo - labour - employment pretty_name: "Discouraged job-seekers by sex, rural / urban area and marital status (thousands) | Asia (ILOSTAT)" --- # Discouraged job-seekers by sex, rural / urban area and marital status (thousands) | Asia (ILOSTAT) 🌏 **4,966 observations** · **25 Asia countries** · **1999–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-4,966-blue) ![countries](https://img.shields.io/badge/countries-25-green) ![years](https://img.shields.io/badge/years-1999–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **4,966 observations** of `Other measures of labour underutilization` data across **25 Asia countries**, spanning **1999–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_WDIS_SEX_GEO_MTS_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Other measures of labour underutilization ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EIP_WDIS_SEX_GEO_MTS_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 25 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `IDN` | 648 | 2000 | 2023 | | `PHL` | 468 | 2007 | 2023 | | `CYP` | 411 | 1999 | 2020 | | `TUR` | 393 | 2000 | 2013 | | `PSE` | 389 | 2012 | 2022 | | `VNM` | 389 | 2010 | 2024 | | `MNG` | 323 | 2013 | 2024 | | `KOR` | 297 | 2015 | 2025 | | `ARM` | 286 | 2008 | 2018 | | `LKA` | 222 | 2016 | 2024 | | `JOR` | 216 | 2017 | 2024 | | `BRN` | 152 | 2014 | 2024 | | `AFG` | 126 | 2014 | 2021 | | `MMR` | 124 | 2015 | 2020 | | `BGD` | 108 | 2010 | 2024 | | ... | _10 more countries_ | | | ## Indicators (sample) - `EIP_WDIS_SEX_GEO_MTS_NB` — Discouraged job-seekers by sex, rural / urban area and marital status (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EIP_WDIS_SEX_GEO_MTS_NB` | | `indicator.label` | `string` | Indicator name in English | `Discouraged job-seekers by sex, rural…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `GEO_COV_NAT` | | `classif1.label` | `string` | — | `Area type: National` | | `classif2` | `string` | Second classification variable where applicable | `MTS_AGGREGATE_TOTAL` | | `classif2.label` | `string` | — | `Marital status (Aggregate): Total` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `135.254` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-eip-wdis-sex-geo-mts-nb-discouraged-job-seekers-by-sex-rural-urban-area-an") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EIP_WDIS_SEX_GEO_MTS_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EIP_WDIS_SEX_GEO_MTS_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EIP_WDIS_SEX_GEO_MTS_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_eip_wdis_sex_geo_mts_nb_discouraged_job_seekers_by_sex_rural_urban_area_an_2025, title = {Discouraged job-seekers by sex, rural / urban area and marital status (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_WDIS_SEX_GEO_MTS_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-eip-wdis-sex-geo-mts-nb-discouraged-job-seekers-by-sex-rural-urban-area-an}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_WDIS_SEX_GEO_MTS_NB_

This dataset contains 4,966 observations of Discouraged job-seekers by sex, rural / urban area and marital status (thousands) data across 25 Asia countries, spanning 1999–2025, covering 1 distinct indicator. The data is sourced from the International Labour Organizations ILOSTAT database, pulled via REST API and filtered to Asia ISO3 country codes. It is structured in tabular format with columns including country code, country name, source, indicator code, indicator name, sex disaggregation, classification variables, year, observed value, observation status, and notes. The indicator is EIP_WDIS_SEX_GEO_MTS_NB, representing discouraged job-seekers disaggregated by sex, rural/urban area, and marital status in thousands. The dataset is suitable for tabular classification, regression, and time-series forecasting tasks, is monolingual (English), has a size category of 1K<n<10K, and is licensed under CC-BY-4.0.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-eip-wdis-sex-geo-mts-nb-discouraged-job-seekers-by-sex-rural-urban-area-an 数据集图片
构建方式
在亚洲劳动力市场研究中,精准捕捉失业群体中的隐蔽失业现象尤为关键。该数据集通过调用国际劳工组织ILOSTAT官方REST API,直接提取了标识符为EIP_WDIS_SEX_GEO_MTS_NB的指标数据,并依据亚洲国家ISO3代码进行地理过滤。来源数据经过ILO统计部门依据国际劳工统计学家会议定义进行统一化处理,使不同国家的调查微观数据得以标准化。数据集中每一观测值均标注了原始来源,确保了数据溯源的可信度与透明度。最终,该数据集由Electric Sheep Asia重新封装为Parquet格式,便于机器学习场景下的直接加载与使用。
特点
该数据集呈现了亚洲劳动力市场中一个独特而重要的维度——因丧失信心而放弃求职的人群,其规模按性别、城乡地域及婚姻状况进行了精细分组。数据集共包含4,966条观测记录,覆盖1999至2025年间的25个亚洲国家,时间跨度极为宽广。除核心指标外,数据还提供了丰富的分类维度,如按性别细分(男性、女性、总计)以及地理覆盖范围(国家、城乡等),为多维度分析提供了坚实基础。此外,数据集中包含观测状态标识与注释信息,便于用户对数据质量与序列中断等潜在问题加以评估。
使用方法
使用者可通过HuggingFace datasets库以一行代码加载该数据集,并利用Pandas快速转换为DataFrame格式进行后续分析。推荐的操作包括:通过'ref_area'列筛选特定国家数据进行纵向研究,或对唯一指标'EIP_WDIS_SEX_GEO_MTS_NB'按时间序列进行可视化探析。借助Pandas的透视表功能,研究者可以轻松将数据重塑为国家×年份矩阵,便于跨国比较与面板数据建模。机器学习应用场景中,此数据集可直接用于回归预测或时间序列预测任务,分类降维后的属性为建模提供了结构化特征集。由于数据以标准化格式存储,无需繁琐的预处理即可接入各类分析流程。
背景与挑战
背景概述
在劳动经济研究中,劳动力利用不足的非标准度量,如沮丧求职者(discouraged job-seekers)的数量,是揭示隐性失业与劳动市场结构性问题的重要指标。该数据集由国际劳工组织(ILO)统计部门于2025年整理发布,并经Electric Sheep Asia封装为机器学习可用格式,聚焦亚洲25个国家1999至2025年间按性别、城乡与婚姻状况细分的沮丧求职者数据。其核心研究问题在于弥补传统失业统计对边缘劳动力群体的覆盖不足,尤其关注亚洲快速城镇化与性别差异下的劳动参与动态。作为ILOSTAT数据库的子集,该数据集为区域劳动力市场比较、政策评估及可持续发展目标(SDG)中的体面劳动监测提供了高质量的标准化基准,推动了跨国劳动经济学的实证分析。
当前挑战
该数据集面临的核心挑战源于所解决的领域问题:沮丧求职者作为隐性失业的代理变量,其定义高度依赖国家劳动调查设计一致性(如ICLS标准),但亚洲各国在调查频率、问卷措辞和分类粒度上存在显著差异,导致跨国可比性受限。在构建过程中,挑战体现为三大技术难题:首先,原始数据源涉及25国多种劳动力调查(如LFS、收入调查等),ILO虽通过‘最佳来源’算法进行协调,但调查体系异质性仍引发断点问题(如方法修订标记‘Break in series’);其次,分类维度(性别、城乡、婚姻状况)的交叉组合导致部分单元格稀疏(如农村女性已婚类别仅有零星观测),影响时间序列建模的稳定性;最后,缺失值管理与观测状态标记(如‘不可靠’标记)的模糊性增加了数据清洗的复杂性,尤其对递归神经网络等序列模型训练构成噪声干扰。
常用场景
经典使用场景
在劳动经济学与人口统计学交叉领域,该数据集为研究亚洲地区隐性失业现象提供了精细化的结构化数据支撑。其核心价值在于揭开了传统失业率统计背后‘沮丧求职者’的阴影群体——那些因长期求职失败而放弃寻找工作的潜在劳动力。研究者可以利用性别、城乡与婚姻状况的交叉分类维度,对亚洲25个国家1999至2025年间的时间序列数据进行多层级剖析,探索经济周期、社会政策与结构性就业障碍如何共同塑造这一群体的规模与演变轨迹。
衍生相关工作
基于该数据集,学界已催生了一系列标志性研究工作。部分学者致力于构建多指标劳动力利用状态矩阵,将沮丧求职者与时间相关就业不足率、潜在劳动力队伍相结合,形成更完整的劳动力闲置指数。另有研究聚焦于婚姻状态如何调节性别与地域对求职意愿的综合效应,运用多水平模型揭示了婚姻承诺对女性重新进入劳动市场的压抑作用。此外,该数据集还推动了动态面板数据方法在劳动经济学中的应用,用以控制未观测异质性对沮丧求职者长期演进趋势的影响。
数据集最近研究
最新研究方向
该数据集聚焦于亚洲地区因性别、城乡地域及婚姻状况分层而产生的失业倦怠群体规模测算,为劳动经济学中隐性失业与劳动力市场边缘化研究提供了关键数据支撑。近期前沿方向集中于利用时间序列与面板数据模型,追踪后疫情时代亚洲劳动力市场中'放弃求职者'的时空演变规律,尤其关注女性、乡村居民等弱势群体在就业信心修复进程中的结构性困境。与此相关的热点事件包括亚洲多国经济复苏背景下,非正规就业膨胀与劳动参与率回升之间的背离现象,该数据有助于揭示官方失业率统计未能捕捉的劳动力市场真实疲软程度,对完善ILO就业质量监测体系、推动包容性增长政策评估具有实证价值。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务