遇见数据集

electricsheepasia/asia-ilo-eip-dwap-sex-age-geo-rt-inactivity-rate-by-sex-age-and-rural-urban-areas

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - asia - ilostat - other-measures-of-labour-underutilization - ilo - labour - employment pretty_name: "Inactivity rate by sex, age and rural / urban areas (%) | Asia (ILOSTAT)" --- # Inactivity rate by sex, age and rural / urban areas (%) | Asia (ILOSTAT) 🌏 **55,244 observations** · **35 Asia countries** · **1970–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-55,244-blue) ![countries](https://img.shields.io/badge/countries-35-green) ![years](https://img.shields.io/badge/years-1970–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **55,244 observations** of `Other measures of labour underutilization` data across **35 Asia countries**, spanning **1970–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_DWAP_SEX_AGE_GEO_RT) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Other measures of labour underutilization ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EIP_DWAP_SEX_AGE_GEO_RT` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 35 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `IDN` | 4,752 | 1990 | 2023 | | `PSE` | 4,416 | 2000 | 2022 | | `CYP` | 3,878 | 1999 | 2025 | | `KHM` | 2,891 | 1996 | 2023 | | `MNG` | 2,736 | 2003 | 2024 | | `ARM` | 2,592 | 2001 | 2023 | | `PHL` | 2,496 | 2007 | 2023 | | `VNM` | 2,496 | 2007 | 2024 | | `PAK` | 2,385 | 2005 | 2025 | | `KOR` | 2,304 | 2000 | 2025 | | `GEO` | 2,304 | 2009 | 2024 | | `THA` | 2,082 | 2007 | 2024 | | `LKA` | 2,034 | 2000 | 2024 | | `IND` | 2,034 | 1994 | 2025 | | `TUR` | 2,016 | 2000 | 2013 | | ... | _20 more countries_ | | | ## Indicators (sample) - `EIP_DWAP_SEX_AGE_GEO_RT` — Inactivity rate by sex, age and rural / urban areas (%) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EIP_DWAP_SEX_AGE_GEO_RT` | | `indicator.label` | `string` | Indicator name in English | `Inactivity rate by sex, age and rural…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHADULT_YGE15` | | `classif1.label` | `string` | — | `Age (Youth, adults): 15+` | | `classif2` | `string` | Second classification variable where applicable | `GEO_COV_NAT` | | `classif2.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `50.27` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `B` | | `obs_status.label` | `string` | — | `Break in series` | | `note_classif` | `string` | — | `C12:2951` | | `note_classif.label` | `string` | — | `Rural / urban areas coverage definiti…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (4 unique values): `SEX_T`, `SEX_M`, `SEX_F`, `SEX_O` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-eip-dwap-sex-age-geo-rt-inactivity-rate-by-sex-age-and-rural-urban-areas") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EIP_DWAP_SEX_AGE_GEO_RT"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EIP_DWAP_SEX_AGE_GEO_RT") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EIP_DWAP_SEX_AGE_GEO_RT"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_eip_dwap_sex_age_geo_rt_inactivity_rate_by_sex_age_and_rural_urban_areas_2025, title = {Inactivity rate by sex, age and rural / urban areas (%) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_DWAP_SEX_AGE_GEO_RT}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-eip-dwap-sex-age-geo-rt-inactivity-rate-by-sex-age-and-rural-urban-areas}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_DWAP_SEX_AGE_GEO_RT_

This dataset contains 55,244 observations of Other measures of labour underutilization data across 35 Asia countries, spanning 1970–2025, covering 1 distinct indicator: Inactivity rate by sex, age and rural / urban areas (%). It is sourced from the International Labour Organization (ILO) ILOSTAT database, repackaged by Electric Sheep Asia, and is designed for tabular classification, regression, and time-series forecasting tasks. The dataset includes variables such as country codes, year, sex, age groups, area types, observed values, and data quality flags, suitable for labour market analysis and machine learning applications.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-eip-dwap-sex-age-geo-rt-inactivity-rate-by-sex-age-and-rural-urban-areas 数据集图片
构建方式
本数据集源自国际劳工组织(ILO)的ILOSTAT中央统计数据库,通过其REST API直接获取原始指标数据,并依据国际劳工统计学家会议(ICLS)定义进行数据协调。数据经过严格筛选,仅保留覆盖亚洲35个国家的观测值,总计55,244条记录,时间跨度从1970年至2025年。数据集以标准化表格形式整理,包含国家代码、性别、年龄、城乡分类、观测值及数据状态等字段,并附有来源标注以便追溯,最终由Electric Sheep Asia团队重新打包发布。
特点
该数据集聚焦于亚洲地区的劳动力未充分利用指标,具体为按性别、年龄及城乡区域划分的不活动率(%)。其显著特点在于多维度的分类变量,允许用户通过性别(如男性、女性、总计)和地域覆盖(国家、城乡)进行精细解析。数据覆盖35个亚洲国家,时间序列长达55年,且标注了数据源类型(如劳动力调查)以及系列中断等质量标记,为用户提供可靠的劳动力市场分析基础。
使用方法
用户可通过HuggingFace的`datasets`库轻松加载数据集:使用`load_dataset`函数即可将数据转换为Pandas DataFrame进行后续分析。典型操作包括按国家筛选数据(如印度尼西亚)、对特定指标进行时间序列可视化,或通过数据透视表构建国家与年份的分析矩阵。数据集适用于分类、回归及时间序列预测任务,支持研究者探索亚洲劳动力市场的长期趋势与结构差异。
背景与挑战
背景概述
劳动力市场的非经济活动率(Inactivity Rate)作为衡量劳动资源未充分利用的关键指标,长期受到国际劳工组织(ILO)及其他发展机构的密切关注。该数据集由国际劳工组织统计部主导,依托其ILOSTAT数据库,经Electric Sheep Asia于2025年重新整理并发布于HuggingFace平台,旨在系统化提供亚洲35个国家自1970年至2025年间按性别、年龄及城乡地域细分的非经济活动率年度观测数据,共计55,244条记录。核心研究问题聚焦于揭示亚洲地区劳动力潜在供给的结构性特征与人口学差异性,为区域劳动力政策制定、劳动经济学建模以及联合国可持续发展目标(SDG)中的体面劳动指标监测提供权威的数据支撑。该数据集因其来源的权威性、跨时域的长跨度以及多维度的分类标准,在区域劳动经济统计领域具有重要的基准参考价值。
当前挑战
该数据集所应对的领域核心挑战在于,非经济活动率作为一个综合性指标,其精准测度高度依赖于不同国家间劳动统计定义的统一性(ICLS标准)及调查方法的可比性,而亚洲各国在经济发展水平、劳动力调查体系及城乡定义上的显著差异,导致跨国家、跨时间序列的纵向比较面临严重的异质性与序列中断问题。在数据集构建过程中,ILOSTAT需对来自35个国家的多种数据来源(如劳动力调查、行政记录)进行繁琐的清洗与同质化处理,并标注数据来源与观察状态(如'序列中断'、'方法修订')以保持透明度;同时,处理多源数据在相同国家与年份下的冲突选择,以及将原始调查的细粒度分类(如性别、城乡)统一纳入标准化模式,构成了主要的技术与协调难点。
常用场景
经典使用场景
在劳动经济学与人口统计学交叉领域,该数据集为分析亚洲地区不同性别、年龄及城乡维度的经济活动人口非参与率提供了标准化、可复现的时序观测基础。经典用法包括:作为时间序列回归模型(如ARIMA、Prophet或GARCH)的训练输入,预测特定国家或区域的非经济活动人口演变轨迹;或作为面板数据固定效应模型的核心因变量,量化城镇化进程、教育扩张及社会保障制度变迁对劳动退出行为的影响幅度。此外,研究者常将其与ILOSTAT其他指标(如失业率、非正规就业比例)联合构建多变量系统,在性别分化视角下剖析亚洲劳动力市场的结构性僵化与韧性。
实际应用
在实际落地中,该数据集支撑着跨国劳动政策的循证评估与国际发展机构的资源配置优化。国际劳工组织与亚洲开发银行可借助这些时序数据监测各国“体面劳动”可持续发展目标(SDG 8)的年度进展,尤其针对女性、青年及农村群体设定减贫与就业促进的量化基准。例如,通过对比城乡区域非经济活动率的季节波动模式,决策者可动态调整农村公共就业计划的启动时机与资金规模。私营部门的人力资源战略规划同样受益:跨国公司可通过目标国家各年龄段的劳动力非参与率趋势,预判未来劳动力供给的紧缺程度,优化工厂选址与自动化投入节奏。
衍生相关工作
基于该数据集已衍生出一系列关联研究与可复现分析管线。国际劳工组织的官方技术报告常以此为基础,构建亚洲各国劳动力市场效率的跨国比较指标体系。在学术领域,研究者利用其多维分类标签(性别、年龄、城乡)训练分类模型,识别高非经济活动风险的亚群体特征,推动了劳动退出预警系统的开发。数据集的时序结构亦催生了若干基准测试任务:在时间序列预测竞赛中,它被用作评估长跨度低频劳动指标预测精度的标准数据源。此外,HuggingFace生态中的社区工作流已将其与天气、基础设施等地理空间数据集拼接,探索气候变迁与劳动力退出的联合动力学模式。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务