遇见数据集

electricsheepasia/asia-ilo-emp-3wap-sex-age-geo-rt-youth-employment-to-population-ratio-by-sex-age-an

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - asia - ilostat - employment - ilo - labour pretty_name: "Youth employment-to-population ratio by sex, age and rural / urban areas (%) | Asia (ILOSTAT)" --- # Youth employment-to-population ratio by sex, age and rural / urban areas (%) | Asia (ILOSTAT) 🌏 **13,131 observations** · **32 Asia countries** · **1970–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-13,131-blue) ![countries](https://img.shields.io/badge/countries-32-green) ![years](https://img.shields.io/badge/years-1970–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **13,131 observations** of `Employment` data across **32 Asia countries**, spanning **1970–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_3WAP_SEX_AGE_GEO_RT) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Employment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EMP_3WAP_SEX_AGE_GEO_RT` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 32 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `IDN` | 1,188 | 1990 | 2023 | | `PSE` | 1,097 | 2000 | 2022 | | `CYP` | 935 | 1999 | 2024 | | `KHM` | 720 | 1996 | 2023 | | `MNG` | 684 | 2003 | 2024 | | `VNM` | 660 | 2006 | 2024 | | `ARM` | 646 | 2001 | 2023 | | `PHL` | 624 | 2007 | 2023 | | `PAK` | 612 | 2005 | 2025 | | `GEO` | 576 | 2009 | 2024 | | `KOR` | 576 | 2000 | 2025 | | `THA` | 516 | 2007 | 2024 | | `LKA` | 504 | 2010 | 2024 | | `TUR` | 504 | 2000 | 2013 | | `IND` | 471 | 1994 | 2025 | | ... | _17 more countries_ | | | ## Indicators (sample) - `EMP_3WAP_SEX_AGE_GEO_RT` — Youth employment-to-population ratio by sex, age and rural / urban areas (%) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EMP_3WAP_SEX_AGE_GEO_RT` | | `indicator.label` | `string` | Indicator name in English | `Youth employment-to-population ratio …` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHBANDS_Y15-29` | | `classif1.label` | `string` | — | `Age (Youth bands): 15-29` | | `classif2` | `string` | Second classification variable where applicable | `GEO_COV_NAT` | | `classif2.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `40.379` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `B` | | `obs_status.label` | `string` | — | `Break in series` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (4 unique values): `SEX_T`, `SEX_M`, `SEX_F`, `SEX_O` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-emp-3wap-sex-age-geo-rt-youth-employment-to-population-ratio-by-sex-age-an") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EMP_3WAP_SEX_AGE_GEO_RT"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EMP_3WAP_SEX_AGE_GEO_RT") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EMP_3WAP_SEX_AGE_GEO_RT"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_emp_3wap_sex_age_geo_rt_youth_employment_to_population_ratio_by_sex_age_an_2025, title = {Youth employment-to-population ratio by sex, age and rural / urban areas (%) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_3WAP_SEX_AGE_GEO_RT}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-emp-3wap-sex-age-geo-rt-youth-employment-to-population-ratio-by-sex-age-an}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_3WAP_SEX_AGE_GEO_RT_

This dataset, titled Youth employment-to-population ratio by sex, age and rural / urban areas (%) | Asia (ILOSTAT), contains tabular data on youth employment-to-population ratios in Asia. Specifically, it includes 13,131 observations across 32 Asian countries, spanning the years 1970 to 2025. The data is sourced from the International Labour Organization (ILO)s ILOSTAT database, a leading global source for labour statistics that harmonizes indicators on employment, unemployment, wages, and more. The dataset features one key indicator: EMP_3WAP_SEX_AGE_GEO_RT, which represents the youth employment-to-population ratio disaggregated by sex, age, and rural/urban areas (%). Data was pulled directly from the ILOSTAT REST API and filtered to Asia ISO3 country codes. The schema includes columns such as country code (ref_area), country name (ref_area.label), source code (source), indicator code (indicator), sex disaggregation (sex), age classification (classif1), area type classification (classif2), year (time), observed value (obs_value), and others, providing detailed breakdowns and metadata. Data quality notes mention that the data is annual, the ILO selects the best source when multiple sources exist, and disaggregation columns are non-null only when published. Repackaged by Electric Sheep Asia, this dataset is part of a mission to offer a unified, ML-ready data layer for Asia on HuggingFace, facilitating easy access for researchers and developers. The license is cc-by-4.0, requiring citation of both the original source and the repackaging.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-emp-3wap-sex-age-geo-rt-youth-employment-to-population-ratio-by-sex-age-an 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT数据库,通过调用其REST API直接获取原始指标数据,并依据亚洲国家ISO3代码进行地理筛选与整合。数据集覆盖32个亚洲经济体,时间跨度为1970年至2025年,包含13,131条观测记录。原始数据经过ILO依据国际劳工统计学家会议(ICLS)定义进行协调处理,确保不同来源(如劳动力调查、住户收入调查等)的数据在概念与分类上的可比性。最终由Electric Sheep Asia重新打包为标准化格式,便于机器学习与时间序列分析场景下的直接调用。
特点
该数据集聚焦于青年就业人口比率这一核心劳动市场指标,按性别、年龄组以及城乡地域进行多维细化。其独特之处在于提供了统一的‘最佳来源’选择逻辑,即对于同一国家与年份存在多个数据来源时,仅保留ILO认定的最优序列,从而减少不一致性。数据模式中包含了丰富的分类维度,如性别(总、男、女、其他)、年龄组(如15-29岁青年带)以及地域覆盖类型(国家或城乡),并附有详细的来源标注与状态标志,便于用户评估数据质量与时间序列中的断点变化。
使用方法
用户可通过HuggingFace Datasets库中的`load_dataset()`函数一步加载该数据,返回的结构化表格可直接转换为pandas DataFrame进行探索。常见分析路径包括:按国家代码筛选特定经济体的子集,绘制特定指标随时间变化的趋势线,或利用透视表操作将数据重塑为国家×年份的面板矩阵。由于数据集已按年度频率组织且包含清晰的分组变量,研究人员可直接用于劳动经济学中的面板数据回归、时间序列预测或分类模型训练,无需额外进行数据清洗与协调。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)于2020年代通过ILOSTAT数据库构建,并由Electric Sheep Asia于2025年重新封装发布,核心研究问题聚焦于亚洲32个国家1970至2025年间青年就业人口比率的性别、年龄及城乡维度差异。作为全球劳动统计的权威来源,ILO依托各国劳动力调查等数据源,依据国际劳动统计学家会议(ICLS)定义进行标准化处理,为评估联合国可持续发展目标(SDG)中体面劳动指标提供了关键数据支撑。该数据集涵盖了超过13,000条观测记录,通过精细的分层维度(性别、年龄组、地理区域)揭示了亚洲青年就业模式的时空演变规律,对发展经济学、劳动经济学及公共政策研究具有重要参考价值。
当前挑战
该数据集所应对的核心领域挑战在于亚洲青年就业比率在性别、年龄及城乡等交叉维度上的巨大异质性——女性、农村及低龄青年往往面临更高的就业参与壁垒,而传统宏观统计难以有效捕捉这种微观差异。构建过程中面临的挑战主要有:1)多国数据源间的口径差异,需通过ICLS标准进行复杂的数据协调与清洗;2)部分国家在时间序列上存在调查方法变更,导致连续性断裂(如数据中的'Break in series'标记);3)不同来源数据的质量参差不齐,需借助ILO选择的'最佳来源'进行优先级判断;4)原始API接口的异构性要求重新封装时进行格式归一化与跨域一致性校验。
常用场景
经典使用场景
该数据集的核心应用场景在于对亚洲青年就业状况进行时序分析与横向比较。研究人员常利用其覆盖32个国家、跨越半个世纪的观测值,构建面板数据模型,以探究青年就业率在性别、年龄段及城乡地域间的演变规律。数据集的按性别与城乡拆分的细致分类,使其成为检验劳动力市场结构性变化、评估宏观经济冲击对青年群体影响的理想素材,亦可服务于国际发展经济学中关于青年就业质量的比较研究。
实际应用
在实际应用中,该数据集为国际发展机构、国家劳动部门及非政府组织制定青年就业政策提供了关键决策依据。通过对不同亚洲国家青年就业率的持续追踪,政策制定者能够识别出那些就业形势严峻的地区与人群——尤其是女性青年与农村青年——进而精准投放职业培训、创业扶持等干预措施。此外,该数据还可辅助评估劳动法规变革、最低工资调整等政策干预的实际效果,为循证决策提供坚实的数据支撑。
衍生相关工作
基于该数据集,研究者已衍生出多项具有影响力的工作。例如,利用面板数据模型,部分研究量化了亚洲国家经济增长与青年就业弹性之间的关系,揭示了不同经济发展阶段青年就业吸纳能力的差异。另有一些工作聚焦于性别分解的数据,构建了亚洲青年就业的性别不平等指数,并探讨了教育普及对缩小就业性别差距的作用。还有学者将本数据集与其他经济发展指标(如人均GDP、城镇化率)进行关联分析,建立了预测模型以评估未来十年亚洲青年就业趋势的变化。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务