遇见数据集

electricsheepasia/asia-ilo-emp-3emp-sex-age-edu-nb-youth-employment-by-sex-age-and-education-thousand

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - asia - ilostat - employment - ilo - labour pretty_name: "Youth employment by sex, age and education (thousands) | Asia (ILOSTAT)" --- # Youth employment by sex, age and education (thousands) | Asia (ILOSTAT) 🌏 **23,480 observations** · **37 Asia countries** · **1970–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-23,480-blue) ![countries](https://img.shields.io/badge/countries-37-green) ![years](https://img.shields.io/badge/years-1970–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **23,480 observations** of `Employment` data across **37 Asia countries**, spanning **1970–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_3EMP_SEX_AGE_EDU_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Employment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EMP_3EMP_SEX_AGE_EDU_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 37 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `IDN` | 1,969 | 1990 | 2023 | | `TUR` | 1,487 | 2000 | 2024 | | `PSE` | 1,393 | 2000 | 2025 | | `KHM` | 1,377 | 1996 | 2023 | | `KOR` | 1,371 | 2000 | 2025 | | `IRN` | 1,364 | 2005 | 2024 | | `CYP` | 1,246 | 1999 | 2024 | | `MNG` | 1,178 | 2000 | 2024 | | `PAK` | 1,037 | 2005 | 2025 | | `THA` | 1,029 | 2000 | 2024 | | `VNM` | 962 | 2010 | 2024 | | `ARM` | 836 | 2001 | 2023 | | `IND` | 810 | 1994 | 2025 | | `LKA` | 786 | 2010 | 2024 | | `GEO` | 780 | 2009 | 2024 | | ... | _22 more countries_ | | | ## Indicators (sample) - `EMP_3EMP_SEX_AGE_EDU_NB` — Youth employment by sex, age and education (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EMP_3EMP_SEX_AGE_EDU_NB` | | `indicator.label` | `string` | Indicator name in English | `Youth employment by sex, age and educ…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHBANDS_Y15-29` | | `classif1.label` | `string` | — | `Age (Youth bands): 15-29` | | `classif2` | `string` | Second classification variable where applicable | `EDU_AGGREGATE_TOTAL` | | `classif2.label` | `string` | — | `Education (Aggregate levels): Total` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `3241.178` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:2620` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (4 unique values): `SEX_T`, `SEX_M`, `SEX_F`, `SEX_O` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-emp-3emp-sex-age-edu-nb-youth-employment-by-sex-age-and-education-thousand") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EMP_3EMP_SEX_AGE_EDU_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EMP_3EMP_SEX_AGE_EDU_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EMP_3EMP_SEX_AGE_EDU_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_emp_3emp_sex_age_edu_nb_youth_employment_by_sex_age_and_education_thousand_2025, title = {Youth employment by sex, age and education (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_3EMP_SEX_AGE_EDU_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-emp-3emp-sex-age-edu-nb-youth-employment-by-sex-age-and-education-thousand}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_3EMP_SEX_AGE_EDU_NB_

This dataset contains youth employment statistics for 37 Asian countries from 1970 to 2025, with 23,480 observations. Sourced from the International Labour Organization (ILO) ILOSTAT database, it covers youth employment by sex, age, and education (in thousands), under the indicator code EMP_3EMP_SEX_AGE_EDU_NB. The dataset includes fields such as country code, year, observed value, data source, sex classification (e.g., total, male, female), age bands, and education levels, suitable for tabular classification, regression, and time-series forecasting tasks. Data is harmonized by ILO using ICLS definitions and includes quality flags (e.g., unreliable, provisional). Repackaged by Electric Sheep Asia for machine learning readiness, it can be loaded directly via the Hugging Face datasets library.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-emp-3emp-sex-age-edu-nb-youth-employment-by-sex-age-and-education-thousand 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT中央统计数据库,通过调用ILOSTAT REST API,提取了标识符为EMP_3EMP_SEX_AGE_EDU_NB的就业指标数据,并依据亚洲ISO3国家代码进行筛选过滤。数据来源涵盖各国劳动力调查、家庭收入调查及行政记录等多元渠道,经ILO依据国际劳动统计学家会议(ICLS)定义进行标准化处理。最终由Electric Sheep Asia团队重新封装为Parquet格式,发布至HuggingFace平台,共收录23,480条观测记录,覆盖37个亚洲国家,时间跨度从1970年至2025年。
特点
该数据集聚焦亚洲地区青年就业的性别、年龄与教育维度,提供了高度细化的分类解析,包含性别(总、男、女、其他)、年龄组(如15-29岁青年段)及教育层级等分类变量。每条观测均附有观测值、状态标记及来源追溯标签,确保了数据的可溯源性。此外,数据集已整合为统一架构,缺失值和多源择优机制经过妥善处理,质量标记如‘不可靠’等状态信息也完整保留,便于用户进行数据筛选与质量评估。
使用方法
用户可通过HuggingFace的datasets库以load_dataset()函数一键加载数据,随后转换为Pandas DataFrame进行灵活操作。典型应用包括按国家代码筛选特定地区的就业数据、针对单一指标(如EMP_3EMP_SEX_AGE_EDU_NB)进行时间序列分析,或通过透视表构建国家×年份的矩阵,以支持跨国比较与趋势研究。数据集结构清晰,包含ref_area, sex, classif1, classif2, time, obs_value等核心字段,便于进行回归、分类及时间序列预测等机器学习任务。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)于2025年通过其核心统计数据库ILOSTAT整理发布,并经Electric Sheep Asia重新包装为机器学习就绪格式。核心研究问题聚焦于亚洲地区青年就业的性别、年龄与教育维度分布,旨在填补对亚洲劳动市场微观结构理解的空白。数据集覆盖37个亚洲国家、1970至2025年间共23,480条观测值,其基于ILO全球劳动力调查与行政记录的标准化处理方法,使其成为研究发展中国家劳动经济学、教育回报率及性别平等议题的重要资源。作为首个面向亚洲青年就业的高分辨率时间序列数据集,它在推动区域劳动政策评估、联合国可持续发展目标(SDG)监测以及计算社会科学领域具有显著影响力。
当前挑战
所解决的领域问题在于劳动统计中普遍存在的数据稀疏性与异质性挑战:亚洲各国在调查频率、分类标准(如教育层级定义)及数据质量上差异显著,传统全球数据集往往牺牲细粒度以换取覆盖度。该数据集通过ILOSTAT的标准化流程(依据国际劳工统计学家会议定义)与质量控制标记(如`obs_status`标识不可靠观测),同时保留原始来源标签以支持溯源性评估,在数据一致性上取得突破。构建过程中面临的挑战包括:①多源数据的时间对齐与缺失值处理,尤其是部分国家早期年份的断档与调查方法变更(由`note_indicator`记录);②性别、年龄与教育三重分类标签的高维稀疏问题,导致某些细分群体的统计数据量不足以支撑可靠推断;③数据集虽标注了37国代码,但如印度尼西亚(1,969行)与阿富汗等国的覆盖深度悬殊,可能引入区域分析中的采样偏差。
常用场景
经典使用场景
该数据集源自国际劳工组织(ILO)的ILOSTAT统计数据库,聚焦于亚洲37个国家1970年至2025年间青年就业状况的细致刻画。其经典使用场景在于通过性别、年龄和教育水平三个核心维度,构建高分辨率的时间序列面板数据,用于分析青年劳动力市场的结构性变迁。研究者常利用此数据集追踪不同性别青年在15至29岁年龄区间内的就业人数与教育层次之间的交互关系,尤其适合开展跨国比较研究,以揭示亚洲各国在青年就业吸纳能力上的异质性模式。
解决学术问题
在学术领域,该数据集精准回应了青年就业研究中的关键问题——如何剥离教育与人口结构对劳动力参与的影响。它解决了传统宏观数据因缺乏细粒度分类而无法深入探讨性别与教育不平等问题的困境,使得学者能够量化分析教育扩张是否有效提升了青年就业率,以及女性青年是否在特定教育层次下面临额外的就业壁垒。该数据集的引入推动了劳动经济学中关于‘教育与就业匹配’假说的实证检验,为理解亚洲新兴经济体的人力资本积累与劳动力市场吸纳能力之间的动态张力提供了可靠的数据基础,对制定精准的青年就业政策具有深远意义。
衍生相关工作
基于该数据集,学术界衍生了若干经典工作方向。例如,利用其多维分类特征,研究者构建了预测青年失业风险的机器学习模型,将性别、年龄与教育水平作为输入特征,探索不同国家情境下就业脆弱群体的画像。另一些工作则聚焦于时间序列分解技术,从长期趋势中分离出周期性波动,以识别亚洲青年就业市场的结构性断点。此外,该数据集还催生了关于教育回报率的跨国比较研究,学者们通过控制性别与年龄变量,重新评估了高等教育在降低青年失业率中的边际效应,为后续将劳动力市场数据与教育政策效果评估相融合提供了方法论范例。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务