遇见数据集

electricsheepasia/asia-ilo-une-tune-sex-edu-geo-nb-unemployment-by-sex-education-and-rural-urban-area

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - asia - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment by sex, education and rural / urban areas (thousands) | Asia (ILOSTAT)" --- # Unemployment by sex, education and rural / urban areas (thousands) | Asia (ILOSTAT) 🌏 **34,725 observations** · **30 Asia countries** · **1970–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-34,725-blue) ![countries](https://img.shields.io/badge/countries-30-green) ![years](https://img.shields.io/badge/years-1970–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **34,725 observations** of `Unemployment` data across **30 Asia countries**, spanning **1970–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_GEO_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_EDU_GEO_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 30 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `IDN` | 3,656 | 1990 | 2023 | | `PSE` | 3,585 | 2000 | 2022 | | `CYP` | 2,675 | 1999 | 2024 | | `MNG` | 2,040 | 2003 | 2024 | | `PAK` | 1,914 | 2005 | 2025 | | `VNM` | 1,899 | 2010 | 2024 | | `KOR` | 1,737 | 2000 | 2025 | | `THA` | 1,703 | 2007 | 2024 | | `IND` | 1,616 | 1994 | 2025 | | `GEO` | 1,548 | 2009 | 2024 | | `ARM` | 1,498 | 2001 | 2023 | | `LKA` | 1,311 | 2010 | 2024 | | `KHM` | 1,228 | 1996 | 2023 | | `JOR` | 975 | 2017 | 2024 | | `BRN` | 928 | 2014 | 2024 | | ... | _15 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_EDU_GEO_NB` — Unemployment by sex, education and rural / urban areas (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_EDU_GEO_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, education and ru…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `classif2` | `string` | Second classification variable where applicable | `GEO_COV_NAT` | | `classif2.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `462.411` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:2620` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (4 unique values): `SEX_T`, `SEX_M`, `SEX_F`, `SEX_O` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-une-tune-sex-edu-geo-nb-unemployment-by-sex-education-and-rural-urban-area") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_EDU_GEO_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_EDU_GEO_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_EDU_GEO_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_une_tune_sex_edu_geo_nb_unemployment_by_sex_education_and_rural_urban_area_2025, title = {Unemployment by sex, education and rural / urban areas (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_GEO_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-une-tune-sex-edu-geo-nb-unemployment-by-sex-education-and-rural-urban-area}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_GEO_NB_

This dataset contains 34,725 observations of unemployment data across 30 Asia countries, spanning from 1970 to 2025, focusing on unemployment disaggregated by sex, education level, and rural/urban areas (in thousands). The data is sourced from the ILOSTAT database of the International Labour Organization (ILO), a leading global source for labour statistics, compiled from national labour force surveys, household income surveys, establishment surveys, and administrative records. The dataset is filtered to Asia ISO3 country codes and harmonized using International Conference of Labour Statisticians (ICLS) definitions. It includes dimensions such as sex (total, male, female, other), education (aggregate total), and area type (national), along with details on data sources, observation status, and notes. The dataset is suitable for tasks like tabular classification, regression, and time-series forecasting, enabling research on labour market trends in Asia.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-une-tune-sex-edu-geo-nb-unemployment-by-sex-education-and-rural-urban-area 数据集图片
构建方式
该数据集以国际劳工组织(ILO)权威统计数据库ILOSTAT为数据源,通过其REST API接口直接抽取指标UNE_TUNE_SEX_EDU_GEO_NB的原始记录,并依据ISO3国家代码筛选出亚洲地区30个经济体的观测数据。ILOSTAT采用国际劳工统计学家会议(ICLS)标准定义对各国劳动力调查、家庭收入调查及行政记录等原始微观数据进行统一调和,数据来源在source.label字段中予以标注,确保了跨国可比性与可追溯性。最终由Electric Sheep Asia进行模式规范化并以Parquet列式格式重新封装发布。
特点
数据集涵盖1970年至2025年间34,725条失业观测记录,覆盖亚洲30个国家,包含性别(男、女、总计及其他)、教育程度与城乡区域等多维分类变量,同时提供观测值、来源标识及数据质量状态标志。所有观测以年度频率呈现,指标单位为千人,缺失分类信息仅在对应维度发布时填充。该数据集结构清晰、字段语义明确,兼具时间序列跨度和横截面异质性,适用于劳动力市场动态分析、性别与教育维度的失业差异研究以及城乡就业结构比较。
使用方法
研究者和开发者可通过HuggingFace datasets库以load_dataset()函数直接加载该数据集,并转换为Pandas DataFrame进行后续分析。典型操作包括按ref_area字段筛选特定国家子集,按indicator字段提取单一指标的时间序列并排序绘图,或利用pivot_table将数据重塑为国家×年份矩阵以开展面板数据分析。该数据集亦可直接用于表格分类、回归及时间序列预测等机器学习任务,使用时应遵循CC-BY-4.0许可协议并同时引用ILO原始来源与Electric Sheep Asia的再封装工作。
背景与挑战
背景概述
伴随全球化进程深化与劳动力市场结构转型,失业问题的多维测度成为劳动经济学与发展研究的核心议题。国际劳工组织(ILO)依托各国劳动力调查与住户收入调查等微观数据,经国际劳工统计学家会议(ICLS)定义协调后,通过ILOSTAT平台发布跨国可比指标。该数据集由Electric Sheep Asia于2025年从ILOSTAT REST API抓取并重新封装,覆盖30个亚洲国家、1970至2025年间34,725条观测,按性别、教育程度与城乡区域三重维度细分失业人数,为亚洲劳动力市场异质性研究提供长时序、高颗粒度的基准数据。
当前挑战
该数据集所应对的领域问题在于,失业率在性别、教育层级与城乡地理之间的结构性差异长期缺乏统一口径下的连续观测,传统宏观总量指标难以揭示弱势群体的就业脆弱性。构建过程中的挑战则源于跨国劳动力调查在抽样设计、教育分类标准与城乡定义上的异质性,ILOSTAT虽以ICLS框架协调,仍存指标断点与方法修订;源数据缺失导致部分国家年份覆盖不均,观测状态标记显示部分数值为暂定或不可靠,且城乡维度仅在国家层面发布时方可获取,限制了细分维度的完整性与跨期可比性。
常用场景
经典使用场景
在劳动经济学与区域发展研究的经典范式中,该数据集通常被用于构建多维度失业率的比较分析框架。研究者依托其性别、教育程度与城乡区划的交叉分类,能够精确刻画亚洲各国在1970至2025年间失业结构的异质性特征。基于时间序列的纵向追踪,可揭示不同教育层级劳动力群体在城乡间的就业脆弱性差异,进而为劳动力市场分层理论提供翔实的经验证据。
衍生相关工作
围绕该数据集,学术界涌现出一系列衍生性研究。比较典型的工作包括基于面板向量自回归模型分析教育扩张对青年失业的时滞效应,以及运用空间计量方法探讨城乡失业率收敛性与区域劳动力市场一体化进程。此外,部分研究将其与ILOSTAT其他指标如工资、工时和非正规就业数据进行链接,构建多维劳动力市场脆弱性指数,拓展了失业分析的深度与广度。
数据集最近研究
最新研究方向
在全球劳动力市场结构性转型与亚洲区域发展不平衡的背景下,该数据集所支撑的前沿研究正朝多维异质性失业动态分析方向演进。依托ILOSTAT对性别、教育程度与城乡区位的交叉分类,研究者得以运用面板时间序列与机器学习方法,精细刻画亚洲各国失业率的时空分异格局。相关热点包括教育与技能错配引致的结构性失业、城乡劳动力市场分割以及女性劳动参与率低迷等议题。该数据为评估教育与就业政策成效、监测可持续发展目标中的体面劳动进展提供了可复现的实证基础,对区域劳动力治理具有重要参考价值。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务