遇见数据集

electricsheepasia/asia-ilo-une-tune-sex-age-mts-nb-unemployment-by-sex-age-and-marital-status-thousan

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 100K<n<1M tags: - tabular - asia - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment by sex, age and marital status (thousands) | Asia (ILOSTAT)" --- # Unemployment by sex, age and marital status (thousands) | Asia (ILOSTAT) 🌏 **123,155 observations** · **35 Asia countries** · **1970–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-123,155-blue) ![countries](https://img.shields.io/badge/countries-35-green) ![years](https://img.shields.io/badge/years-1970–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **123,155 observations** of `Unemployment` data across **35 Asia countries**, spanning **1970–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_MTS_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_AGE_MTS_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 35 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `IDN` | 9,457 | 1996 | 2023 | | `KOR` | 9,050 | 2000 | 2025 | | `TUR` | 8,806 | 2000 | 2024 | | `PHL` | 7,940 | 2001 | 2023 | | `IRN` | 6,721 | 2005 | 2024 | | `CYP` | 6,155 | 1999 | 2020 | | `ARM` | 6,049 | 2001 | 2023 | | `VNM` | 5,239 | 2010 | 2024 | | `MNG` | 5,003 | 2009 | 2024 | | `THA` | 4,980 | 2000 | 2024 | | `ISR` | 4,648 | 2012 | 2024 | | `PAK` | 4,547 | 2005 | 2025 | | `JPN` | 4,481 | 2000 | 2023 | | `IND` | 4,321 | 1994 | 2025 | | `LKA` | 3,932 | 2010 | 2024 | | ... | _20 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_AGE_MTS_NB` — Unemployment by sex, age and marital status (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_AGE_MTS_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, age and marital …` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHADULT_YGE15` | | `classif1.label` | `string` | — | `Age (Youth, adults): 15+` | | `classif2` | `string` | Second classification variable where applicable | `MTS_AGGREGATE_TOTAL` | | `classif2.label` | `string` | — | `Marital status (Aggregate): Total` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `462.411` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `—` | | `note_classif.label` | `string` | — | `—` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (4 unique values): `SEX_T`, `SEX_M`, `SEX_F`, `SEX_O` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-une-tune-sex-age-mts-nb-unemployment-by-sex-age-and-marital-status-thousan") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_MTS_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_AGE_MTS_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_MTS_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_une_tune_sex_age_mts_nb_unemployment_by_sex_age_and_marital_status_thousan_2025, title = {Unemployment by sex, age and marital status (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_MTS_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-une-tune-sex-age-mts-nb-unemployment-by-sex-age-and-marital-status-thousan}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_MTS_NB_

This dataset contains 123,155 observations of unemployment data (in thousands) by sex, age, and marital status across 35 Asia countries, spanning from 1970 to 2025. The data is sourced from the ILOSTAT database of the International Labour Organization (ILO), retrieved via API and filtered to Asian countries, covering the indicator UNE_TUNE_SEX_AGE_MTS_NB. It includes detailed columns such as country codes, sex disaggregation (total, male, female, etc.), age and marital status classifications, observation year, unemployment values, and data quality flags (e.g., reliability status). The data is annual in frequency and harmonized by ILO standards, suitable for tabular classification, regression, and time-series forecasting tasks.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-une-tune-sex-age-mts-nb-unemployment-by-sex-age-and-marital-status-thousan 数据集图片
构建方式
在全球劳动力市场统计体系中,国际劳工组织(ILO)长期承担着跨国就业与失业数据的标准化职责,ILOSTAT数据库便是这一职能的核心载体。本数据集经由Electric Sheep Asia团队从ILOSTAT REST API接口直接抽取指标UNE_TUNE_SEX_AGE_MTS_NB的原始记录,并依据ISO 3166-1 alpha-3国家代码筛选出亚洲区域35个国家,最终形成123,155条观测。数据涵盖1970年至2025年的年度序列,ILO统计部门依照国际劳工统计学家会议(ICLS)的定义对各国劳动力调查微观数据进行统一协调,来源信息在source.label列中予以标注,以保障数据溯源的可查性。
特点
该数据集以亚洲地区失业人口规模为核心测度,按性别、年龄与婚姻状况三重维度进行交叉分类,涵盖SEX_T、SEX_M、SEX_F、SEX_O四类性别标识及多种年龄与婚姻状态分组。时序跨度逾半个世纪,最早可追溯至1970年,最晚延伸至2025年,为长时段劳动力市场变迁研究提供了罕有的连续观测基础。数据结构采用扁平化表格形式,包含国家代码、来源编码、指标代码、分类变量及其标签、观测年份、观测值、观测状态标志与注释信息等字段,兼具可读性与机器可处理性,适用于分类、回归及时间序列预测等多类任务。
使用方法
研究者可通过HuggingFace datasets库以load_dataset()函数直接加载该数据集,并借助to_pandas()方法转换为DataFrame进行后续分析。典型操作包括按ref_area列筛选特定国家子集,例如提取印度尼西亚(IDN)的全部记录;亦可针对单一指标按时间排序后绘制趋势图,以观察失业规模的历年演变。若需开展跨国比较研究,可利用pivot_table()方法将数据重塑为国家×年份的矩阵形式,便于横向对照。此外,观测状态标志(obs_status)与注释字段(note_indicator、note_source)可用于识别数据质量异常或序列中断情形,为分析中的稳健性检验提供依据。
背景与挑战
背景概述
国际劳工组织(ILO)自1919年成立以来,始终致力于全球劳动统计的标准化与传播,其ILOSTAT数据库堪称劳动领域最权威的跨国数据来源。在此背景下,Electric Sheep Asia于2025年对ILOSTAT原始数据进行再包装,发布了涵盖35个亚洲国家、1970至2025年间123,155条观测的失业数据集,按性别、年龄与婚姻状况细分。该数据集的核心研究问题在于揭示亚洲劳动力市场中失业的结构性差异,为劳动经济学、社会政策评估及可持续发展目标(SDG)监测提供实证基础。其影响力在于以机器学习就绪的格式降低数据获取门槛,推动亚洲劳动市场的量化研究与跨区域比较。
当前挑战
该数据集所解决的领域问题在于精准刻画亚洲失业现象的多元维度:通过性别、年龄与婚姻状况的交叉分类,揭示不同人口群体的失业风险差异,为制定针对性就业政策提供依据。其构建过程中的挑战尤为显著:原始数据源自各国劳动调查,在ILO协调下面临定义不一致、报告周期错位、部分年份数据缺失或标记为“不可靠”等难题;数据集需处理来源异质性、断点修订(如方法论变更)以及分类变量非全指标覆盖等问题,同时需在保留原始溯源信息与实现机器学习友好格式之间取得平衡。这些挑战共同构成了数据质量与可用性的核心张力。
常用场景
经典使用场景
在劳动经济学与人口统计学的交叉领域,失业率的性别、年龄与婚姻状况分异历来是洞察劳动力市场结构性特征的关键切口。该数据集汇聚亚洲35国自1970年至2025年间12万余条观测,按性别、年龄组与婚姻状态三重维度细粒度记录失业规模,其最经典的使用场景即在于支撑跨国、跨时期的失业异质性比较研究,例如探究不同婚姻状态群体在经济周期中的脆弱性差异,或分析青年与成年失业率的长期收敛趋势。
衍生相关工作
以该数据集为基石,衍生出一系列围绕亚洲失业动态的经典研究,包括跨国失业收敛性的面板计量分析、婚姻状态与失业持续期的生存模型构建,以及基于机器学习方法的失业率短期预测。这些工作广泛发表于劳动经济学期刊与国际劳工组织的旗舰报告中,进一步推动了ILOSTAT数据库在学术与政策圈层的深度应用,并催生了针对亚洲特定次区域失业模式的专题比较研究。
数据集最近研究
最新研究方向
在全球劳动力市场性别差距与人口老龄化议题持续升温的背景下,该数据集凭借其覆盖35个亚洲国家、跨越1970至2025年的12.3万余条观测,为失业率的异质性分析提供了关键支撑。当前前沿研究借助该数据集探究性别、年龄与婚姻状况交互作用下失业风险的动态演化,尤其关注青年与女性群体在婚姻状态转变中的脆弱性。相关热点事件如国际劳工组织对非正规就业与SDG体面劳动目标的追踪,进一步凸显了该数据在政策评估与时间序列预测中的价值,为亚洲地区劳动力市场韧性研究奠定了实证基础。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务