遇见数据集

electricsheepasia/asia-ilo-une-tune-sex-age-cbr-nb-unemployment-by-sex-age-and-place-of-birth-thousan

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 1K<n<10K tags: - tabular - asia - ilostat - international-migrant-stock - ilo - labour - employment pretty_name: "Unemployment by sex, age and place of birth (thousands) | Asia (ILOSTAT)" --- # Unemployment by sex, age and place of birth (thousands) | Asia (ILOSTAT) 🌏 **9,242 observations** · **27 Asia countries** · **1991–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-9,242-blue) ![countries](https://img.shields.io/badge/countries-27-green) ![years](https://img.shields.io/badge/years-1991–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **9,242 observations** of `International migrant stock` data across **27 Asia countries**, spanning **1991–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_CBR_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** International migrant stock ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_AGE_CBR_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 27 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `CYP` | 1,950 | 1999 | 2025 | | `TUR` | 1,401 | 2000 | 2025 | | `ISR` | 1,053 | 2012 | 2024 | | `ARM` | 772 | 2001 | 2024 | | `IDN` | 559 | 2017 | 2024 | | `BRN` | 524 | 2014 | 2024 | | `TLS` | 322 | 2001 | 2022 | | `THA` | 319 | 2018 | 2024 | | `ARE` | 300 | 2018 | 2024 | | `MNG` | 250 | 2019 | 2024 | | `MDV` | 208 | 2014 | 2019 | | `KHM` | 208 | 2008 | 2021 | | `IRN` | 202 | 2006 | 2011 | | `IRQ` | 198 | 2007 | 2021 | | `OMN` | 189 | 2021 | 2022 | | ... | _12 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_AGE_CBR_NB` — Unemployment by sex, age and place of birth (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:6361` | | `source.label` | `string` | Source name in English | `HIES - Households Living Conditions S…` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_AGE_CBR_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, age and place of…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHADULT_YGE15` | | `classif1.label` | `string` | — | `Age (Youth, adults): 15+` | | `classif2` | `string` | Second classification variable where applicable | `CBR_BIR_TOTAL` | | `classif2.label` | `string` | — | `Place of birth: Total` | | `time` | `int64` | Observation year | `2014` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `601.892` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `—` | | `note_classif.label` | `string` | — | `—` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-une-tune-sex-age-cbr-nb-unemployment-by-sex-age-and-place-of-birth-thousan") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_CBR_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_AGE_CBR_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_CBR_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_une_tune_sex_age_cbr_nb_unemployment_by_sex_age_and_place_of_birth_thousan_2025, title = {Unemployment by sex, age and place of birth (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_CBR_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-une-tune-sex-age-cbr-nb-unemployment-by-sex-age-and-place-of-birth-thousan}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_CBR_NB_

This dataset contains unemployment statistics for 27 Asian countries from 1991 to 2025, with 9,242 observations, covering one core indicator: Unemployment by sex, age and place of birth (thousands). The data is sourced from the International Labour Organization (ILO) ILOSTAT database, retrieved via API and filtered to Asian countries. It provides detailed disaggregation dimensions, including sex (total, male, female), age groups (e.g., 15+), and place of birth (total), along with country codes, year, observed values, data sources, and quality flags. The data is published at annual frequency and is suitable for tasks such as tabular classification, regression, and time-series forecasting.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-une-tune-sex-age-cbr-nb-unemployment-by-sex-age-and-place-of-birth-thousan 数据集图片
构建方式
该数据集以国际劳工组织(ILO)的ILOSTAT中央统计数据库为数据源,通过ILOSTAT REST API直接抽取指标UNE_TUNE_SEX_AGE_CBR_NB的原始记录,并依据ISO3国家代码筛选出27个亚洲国家。ILOSTAT依据国际劳工统计学家会议(ICLS)定义对各国劳动力调查、家庭收入调查等微观数据进行调和处理,数据集在重打包过程中保留source.label等溯源字段,确保数据来源可追溯,最终形成覆盖1991至2025年的标准化表格数据。
特点
数据集包含9,242条观测记录,覆盖27个亚洲国家,时间跨度自1991年至2025年,以年度频率发布。核心字段涵盖国家代码、数据来源、指标编码、性别、年龄分类、出生地分类、年份及观测值,并附有观测状态标记与注释信息。性别维度提供总计、男性、女性三类取值,年龄与出生地分类按指标发布情况动态填充,适用于分类、回归及时间序列预测等多类表格任务。
使用方法
研究者可通过HuggingFace datasets库以load_dataset函数加载数据集,并转换为Pandas DataFrame进行后续分析。典型操作包括按ref_area字段筛选特定国家、按indicator字段选取单一指标并依时间排序绘制趋势图,以及利用pivot_table将数据重塑为国家与年份交叉的矩阵形式。该数据集亦可用于跨国失业率的横向比较、性别与年龄维度的差异化分析以及时间序列建模。
背景与挑战
背景概述
在全球劳动力市场日益关注移民就业融入的背景下,国际劳工组织(ILO)依托其权威统计数据库ILOSTAT,构建了按性别、年龄及出生地分类的失业数据体系。该数据集由Electric Sheep Asia于2026年重新封装发布,覆盖27个亚洲国家、1991至2025年间共计9242条观测记录,核心指标聚焦于移民与本土出生人口的失业差异。作为ILOSTAT官方数据的标准化镜像,它继承了国际劳工统计学家会议(ICLS)的定义框架,为探究亚洲区域内移民劳动力市场边缘化问题提供了跨时空、可复现的基准数据,对劳动经济学与移民政策研究具有基础性支撑价值。
当前挑战
该数据集所回应的领域难题在于,移民失业状况的跨国可比性长期受制于各国统计口径、调查方法与劳动力定义的分歧,难以支撑可靠的比较分析。构建过程中,源数据由多个国家劳动力调查与住户收入调查汇编而成,部分国家时序严重缺失或年份断档,且大量观测值带有“不可靠”或“方法修订导致序列断裂”等质量标识。此外,性别、年龄与出生地维度的交叉分类仅在指标发布该细分时才非空,致使数据集内部存在结构性稀疏,对模型训练中的缺失值处理与偏差校正构成显著挑战。
常用场景
经典使用场景
在劳动经济学与人口迁移研究的交叉领域,该数据集最为经典的运用在于按性别、年龄组与出生地维度对亚洲27国的失业规模进行跨国比较与历时追踪。研究者可借由1991至2025年的年度观测值,构建面板数据模型,考察本土出生与外国出生劳动力在失业风险上的结构性差异,亦可借助时间序列分解方法识别不同年龄队列失业波动的周期特征。
实际应用
在政策实践层面,该数据集可服务于亚洲各国劳工部门与移民管理机构,用于监测外国出生劳动力群体的失业动态并识别高风险人群,从而为就业服务配置、职业培训计划设计与社会保障覆盖扩展提供量化依据。国际组织亦可利用其编制区域移民就业脆弱性指数,辅助制定面向特定性别与年龄群体的精准干预方案,提升劳动力市场治理的针对性与时效性。
衍生相关工作
围绕该数据集及其源指标,已衍生出一系列经典性工作,包括ILO定期发布的《国际移民劳动者估计》报告、全球移民数据门户的跨国比较分析,以及学术界针对海湾合作委员会国家与东南亚经济体移民失业问题的实证论文。Electric Sheep Asia的Parquet重打包版本进一步降低了机器学习社区的使用门槛,催生了以该数据为基础的表格回归与时间序列预测基准实验,拓展了劳动统计数据的再利用边界。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务