遇见数据集

electricsheepasia/asia-ilo-eip-teip-sex-geo-nb-persons-outside-the-labour-force-by-sex-and-rural

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 1K<n<10K tags: - tabular - asia - ilostat - other-measures-of-labour-underutilization - ilo - labour - employment pretty_name: "Persons outside the labour force by sex and rural / urban areas (thousands) | Asia (ILOSTAT)" --- # Persons outside the labour force by sex and rural / urban areas (thousands) | Asia (ILOSTAT) 🌏 **3,222 observations** · **31 Asia countries** · **1970–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-3,222-blue) ![countries](https://img.shields.io/badge/countries-31-green) ![years](https://img.shields.io/badge/years-1970–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **3,222 observations** of `Other measures of labour underutilization` data across **31 Asia countries**, spanning **1970–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_TEIP_SEX_GEO_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Other measures of labour underutilization ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EIP_TEIP_SEX_GEO_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 31 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `IDN` | 297 | 1990 | 2023 | | `PSE` | 276 | 2000 | 2022 | | `CYP` | 234 | 1999 | 2024 | | `KHM` | 180 | 1996 | 2023 | | `MNG` | 171 | 2003 | 2024 | | `ARM` | 162 | 2001 | 2023 | | `VNM` | 156 | 2007 | 2024 | | `PHL` | 156 | 2007 | 2023 | | `PAK` | 153 | 2005 | 2025 | | `GEO` | 144 | 2009 | 2024 | | `KOR` | 144 | 2000 | 2025 | | `IND` | 138 | 1994 | 2025 | | `THA` | 129 | 2007 | 2024 | | `LKA` | 126 | 2010 | 2024 | | `TUR` | 126 | 2000 | 2013 | | ... | _16 more countries_ | | | ## Indicators (sample) - `EIP_TEIP_SEX_GEO_NB` — Persons outside the labour force by sex and rural / urban areas (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EIP_TEIP_SEX_GEO_NB` | | `indicator.label` | `string` | Indicator name in English | `Persons outside the labour force by s…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `GEO_COV_NAT` | | `classif1.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `8230.246` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `B` | | `obs_status.label` | `string` | — | `Break in series` | | `note_classif` | `string` | — | `—` | | `note_classif.label` | `string` | — | `—` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (4 unique values): `SEX_T`, `SEX_M`, `SEX_F`, `SEX_O` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-eip-teip-sex-geo-nb-persons-outside-the-labour-force-by-sex-and-rural") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EIP_TEIP_SEX_GEO_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EIP_TEIP_SEX_GEO_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EIP_TEIP_SEX_GEO_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_eip_teip_sex_geo_nb_persons_outside_the_labour_force_by_sex_and_rural_2025, title = {Persons outside the labour force by sex and rural / urban areas (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_TEIP_SEX_GEO_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-eip-teip-sex-geo-nb-persons-outside-the-labour-force-by-sex-and-rural}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EIP_TEIP_SEX_GEO_NB_

This dataset contains 3,222 observations of Other measures of labour underutilization data across 31 Asia countries, spanning from 1970 to 2025, covering 1 distinct indicator: EIP_TEIP_SEX_GEO_NB, which represents persons outside the labour force by sex and rural/urban areas in thousands. The data is sourced from the International Labour Organizations (ILO) ILOSTAT database, retrieved via API and filtered to Asian countries, providing annual labour statistics. It includes columns such as country codes, sources, indicators, sex disaggregation, time years, and observed values, suitable for tabular classification, regression, and time-series forecasting tasks. The dataset is normalized for machine learning applications and released under the CC-BY-4.0 license.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-eip-teip-sex-geo-nb-persons-outside-the-labour-force-by-sex-and-rural 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT核心统计数据库,通过REST API接口直接抓取指标代码为EIP_TEIP_SEX_GEO_NB的原始数据,并依据亚洲ISO3国家代码进行地理筛选。ILOSTAT依据国际劳工统计学家会议(ICLS)定义对各国劳动力调查的微观数据进行统一协调,数据源类型在source.label列中予以标注,确保了跨国的可比性与可追溯性。最终由Electric Sheep Asia重新打包,形成包含3222条观测值、覆盖31个亚洲国家、时间跨度为1970年至2025年的标准化表格数据集。
使用方法
数据集可直接通过HuggingFace Datasets库加载,调用load_dataset函数即可获得pandas DataFrame,便于进行数据探索与建模。用户可按国家代码过滤单一国家的时间序列,或针对特定指标按年份排序后绘制趋势图。亦可通过透视表操作,将数据重塑为国家×年份的矩阵形式,以进行跨国比较或面板回归分析。数据集兼容表格分类、回归以及时间序列预测等多种任务,适用于劳动经济学、社会政策研究等领域的研究者与数据科学家。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)于2025年通过Electric Sheep Asia在HuggingFace平台重新封装发布,核心研究问题聚焦于亚洲31个国家中因性别与城乡地理差异而游离于劳动力市场之外的群体规模。数据集整合了ILOSTAT数据库自1970年至2025年间共计3,222条观测记录,其指标EIP_TEIP_SEX_GEO_NB为理解亚洲地区劳动力利用不足的复杂结构提供了关键的定量基础。作为ILO中央统计数据仓库的分支,该数据集不仅继承了官方统计的权威性,更通过标准化的细分维度(性别与城乡)揭示了非劳动力人口在宏观经济与社会政策评估中的隐性贡献。其发布对劳动经济学、发展社会学以及国际比较研究领域产生了深远影响,为衡量性别平等和城乡发展差距等可持续发展目标提供了精细化的数据支撑。
当前挑战
该数据集所解决的领域核心问题在于,传统劳动力统计往往忽略或简化劳动力市场外围群体——即不活跃人口——的内部异质性。性别与城乡二元交叉分析揭示了被宏观数据掩盖的静态与动态非参与模式,例如女性因家庭照料责任而退出市场、或是农村地区因产业结构单一导致的隐性失业。在构建过程中,挑战主要源于各国调查方法论的差异与时间序列的断裂,如数据集中标注的“Break in series”和“Methodology revised”反映了ILO在协调国家劳动力调查(LFS)、家庭收支调查等多种原始数据源时遭遇的一致性难题。此外,部分国家历史数据稀疏(如阿富汗仅有有限记录),以及年度数据与高频月度/季度数据之间的粒度断层,均限制了模型对短期波动与突发事件(如经济危机)冲击效应的捕捉能力。
常用场景
经典使用场景
在劳动经济学与人口统计学的交叉研究中,该数据集被广泛用于分析亚洲地区劳动力市场边缘群体的构成与演变。其核心价值在于提供了按性别和城乡地理区域细分的非劳动力人口时序数据,覆盖31个亚洲国家长达半个世纪的观测。研究者可借此构建面板数据模型,探讨经济增长、城镇化进程与劳动力退出之间的关联机制,或评估社会保障政策对弱势群体的覆盖效果。
解决学术问题
该数据集精准填补了亚洲劳动经济学研究中关于'潜在劳动力'群体量化分析的空白。传统失业率指标无法反映因家庭责任、教育限制或疾病等原因退出劳动市场的隐性失业问题。通过引入ILO标准化定义的非劳动力人口分类,研究者得以量化性别差异与城乡差距在劳动力市场边缘化过程中的复合效应,为理解亚洲特有的非正规就业与劳动力资源错配现象提供了可靠的实证基础。
实际应用
在政策制定与评估领域,该数据集被国际劳工组织及各国统计部门用于监测可持续发展目标中体面劳动指标的进展。具体而言,决策者可依据城乡与性别维度的非劳动力人口变化趋势,精准识别就业服务、技能培训或社会救助政策的优先干预区域。例如,若某地区农村女性非劳动力人口持续上升,则暗示需加强该群体的就业支持或社保覆盖,从而优化公共资源的配置效率。
数据集最近研究
最新研究方向
当前,国际劳工组织(ILO)发布的亚洲地区劳动力市场数据集聚焦于性别与城乡维度下‘劳动力市场以外人员’的统计,这一方向与全球关注的非正规就业、劳动参与率下降及可持续发展目标(SDG 8)紧密关联。该数据集覆盖31个亚洲国家长达55年的时序观测,为研究经济结构转型中劳动力边缘化现象提供了稀缺的标准化基准。前沿研究正借助此类高颗粒度数据,结合空间计量与时序预测模型,剖析新冠疫情后亚洲女性劳动退出率的城乡差异,以及城镇化进程中隐性失业的时空演变规律。其意义在于,通过量化分析促进政策制定者识别脆弱群体,优化社会保护体系,同时推动劳动统计与机器学习交叉领域的方法论创新,为亚太地区的包容性经济增长提供数据驱动的决策支持。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务