遇见数据集

electricsheepasia/asia-ilo-une-tune-sex-edu-dsb-nb-unemployment-by-sex-education-and-disability-statu

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 1K<n<10K tags: - tabular - asia - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment by sex, education and disability status (thousands) | Asia (ILOSTAT)" --- # Unemployment by sex, education and disability status (thousands) | Asia (ILOSTAT) 🌏 **4,590 observations** · **20 Asia countries** · **1996–2024** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-4,590-blue) ![countries](https://img.shields.io/badge/countries-20-green) ![years](https://img.shields.io/badge/years-1996–2024-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **4,590 observations** of `Unemployment` data across **20 Asia countries**, spanning **1996–2024**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_DSB_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_EDU_DSB_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 20 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `CYP` | 661 | 2005 | 2024 | | `MNG` | 652 | 2006 | 2024 | | `ARM` | 559 | 2007 | 2023 | | `ISR` | 354 | 2016 | 2023 | | `IDN` | 348 | 2010 | 2023 | | `KHM` | 293 | 1996 | 2023 | | `LKA` | 233 | 2018 | 2024 | | `PSE` | 210 | 2018 | 2022 | | `BGD` | 193 | 2011 | 2024 | | `THA` | 179 | 2007 | 2019 | | `TLS` | 156 | 2015 | 2022 | | `IRQ` | 119 | 2007 | 2021 | | `AFG` | 118 | 2017 | 2021 | | `LAO` | 110 | 2015 | 2022 | | `TUR` | 93 | 2000 | 2024 | | ... | _5 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_EDU_DSB_NB` — Unemployment by sex, education and disability status (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_EDU_DSB_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, education and di…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `classif2` | `string` | Second classification variable where applicable | `DSB_STATUS_TOTAL` | | `classif2.label` | `string` | — | `Disability status: Total` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `462.411` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:2620` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-une-tune-sex-edu-dsb-nb-unemployment-by-sex-education-and-disability-statu") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_EDU_DSB_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_EDU_DSB_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_EDU_DSB_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_une_tune_sex_edu_dsb_nb_unemployment_by_sex_education_and_disability_statu_2024, title = {Unemployment by sex, education and disability status (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2024}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_DSB_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-une-tune-sex-edu-dsb-nb-unemployment-by-sex-education-and-disability-statu}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_DSB_NB_

This dataset contains unemployment data for 20 Asian countries, disaggregated by sex, education level, and disability status, with values in thousands. It spans from 1996 to 2024, comprising 4,590 observations and one core indicator (UNE_TUNE_SEX_EDU_DSB_NB). The data is sourced from the International Labour Organization (ILO) ILOSTAT database, retrieved via API and filtered to Asian ISO3 country codes. The dataset includes fields such as country code, country name, data source, indicator code, sex classification, education classification, disability status classification, year, observed value, and observation status, supporting tabular classification, regression, and time-series forecasting tasks. The data is annual frequency, with ILO selecting the best source for each country-year combination, and disaggregation columns are non-null only when the indicator publishes that breakdown. It is licensed under CC-BY-4.0 and repackaged by Electric Sheep Asia for machine learning readiness.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-une-tune-sex-edu-dsb-nb-unemployment-by-sex-education-and-disability-statu 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的中央统计数据库ILOSTAT,通过其REST API接口直接提取指标UNE_TUNE_SEX_EDU_DSB_NB的原始记录,并依据ISO3国家代码筛选出亚洲地区20个国家的数据。原始调查微数据经ILO统计局依照国际劳工统计学家会议(ICLS)定义进行标准化处理,同时保留来源标签以确保可追溯性。最终由Electric Sheep Asia重新打包为Parquet格式,并发布至HuggingFace平台,形成覆盖1996至2024年、共计4590条观测值的结构化表格数据集。
特点
该数据集以亚洲20国为地理范围,时间跨度长达28年,聚焦于按性别、教育程度和残疾状况分列的失业人口(千人)统计。其核心特点在于多维度分类编码的完整性,包括性别(总计、男性、女性)、教育水平聚合类别及残疾状况状态,并附有观测状态标志和来源注释。数据以年度频率呈现,每条记录均包含国家代码、指标代码、分类变量及数值,便于进行细粒度的横截面与时间序列分析。此外,数据集遵循CC-BY-4.0许可,兼具权威性与开放性。
使用方法
研究者可通过HuggingFace的datasets库以一行代码加载数据集,并转换为Pandas数据框进行灵活操作。典型用法包括按国家代码筛选特定经济体的失业记录,或针对单一指标绘制时间序列趋势图。用户亦可利用透视表功能将数据重塑为国家与年份交叉的矩阵,以支持面板数据分析。在分类或回归任务中,性别、教育程度和残疾状况等分类变量可作为特征输入,而观测值则作为预测目标。使用时需注意年度频率限制及部分观测状态标志所提示的数据质量差异。
背景与挑战
背景概述
国际劳工组织(ILO)长期致力于全球劳动力市场统计监测,其ILOSTAT数据库为劳动力研究提供权威数据支撑。该数据集由Electric Sheep Asia于2026年重新封装,源自ILO的UNE_TUNE_SEX_EDU_DSB_NB指标,涵盖1996至2024年间20个亚洲国家的4590条失业观测记录,按性别、教育程度和残疾状况多维细分。数据集旨在回应亚洲地区失业问题的异质性分析需求,为探究教育水平、性别差异与残疾状态对就业机会的影响提供量化基础,对劳动力经济学与包容性就业政策研究具有重要参考价值。
当前挑战
该数据集所对应的领域问题聚焦于多维失业率的精准刻画与跨国可比性。构建过程中面临多重挑战:各国劳动力调查在抽样设计、教育分类标准及残疾定义上存在显著异质性,ILO虽依据国际劳工统计学家会议(ICLS)定义进行协调,但数据缺失、序列断裂及不可靠标志仍普遍存在;部分国家观测年份不连续,残疾状况维度的数据覆盖尤为稀疏,导致细粒度交叉分析受限;此外,非标准教育分类与来源更替引发的统计口径变化,对时序建模和跨国回归的稳健性构成持续挑战。
常用场景
经典使用场景
在劳动经济学与不平等研究的经典图景中,该数据集常被用于构建亚洲地区失业率的性别—教育—残疾三重交叉分析框架。研究者借助其长时段面板结构,考察不同受教育程度与残疾状态下男女失业率的差异化演变轨迹,并通过时间序列建模揭示经济周期中弱势群体的就业脆弱性。其按ISO3国家代码与年度观测值组织的表格形态,亦便于开展跨国比较与面板回归分析。
衍生相关工作
围绕该数据集已衍生出若干经典研究方向。部分学者将其与ILOSTAT其他指标合并,构建亚洲劳动力市场脆弱性指数;亦有研究以此为基础,运用机器学习方法对失业率进行短期预测与情景模拟。在比较政治经济学领域,该数据被用于检验教育扩张对性别就业差距的收敛效应,以及残疾包容政策与失业率变动之间的因果关联,形成了跨学科的方法论积累。
数据集最近研究
最新研究方向
在全球劳动力市场结构性变革与包容性增长议程交织的背景下,该数据集所承载的性别、教育程度与残疾状况三重维度的失业细分信息,正成为劳动经济学与残障研究交叉领域的前沿切入点。当前研究愈发关注教育禀赋如何调节残疾群体的就业脆弱性,以及性别差异在其中扮演的叠加角色,尤其是在亚洲新兴经济体非正规就业盛行的语境下。ILO近期推动的残疾包容性就业指标框架与SDG第8项体面工作目标,使此类数据成为监测结构性排斥、评估技能匹配政策成效的关键证据基础。相关成果对优化社会保障靶向机制、弥合弱势群体就业鸿沟具有实证支撑意义。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务