遇见数据集

electricsheepasia/asia-ilo-ees-tees-sex-geo-nb-employees-by-sex-and-rural-urban-areas-thousands

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 1K<n<10K tags: - tabular - asia - ilostat - employees - ilo - labour - employment pretty_name: "Employees by sex and rural / urban areas (thousands) | Asia (ILOSTAT)" --- # Employees by sex and rural / urban areas (thousands) | Asia (ILOSTAT) 🌏 **3,132 observations** · **30 Asia countries** · **1970–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-3,132-blue) ![countries](https://img.shields.io/badge/countries-30-green) ![years](https://img.shields.io/badge/years-1970–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **3,132 observations** of `Employees` data across **30 Asia countries**, spanning **1970–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_GEO_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Employees ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EES_TEES_SEX_GEO_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 30 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `PSE` | 276 | 2000 | 2022 | | `IDN` | 252 | 1996 | 2023 | | `CYP` | 234 | 1999 | 2024 | | `KHM` | 180 | 1996 | 2023 | | `MNG` | 171 | 2003 | 2024 | | `ARM` | 162 | 2001 | 2023 | | `PHL` | 156 | 2007 | 2023 | | `VNM` | 156 | 2007 | 2024 | | `PAK` | 153 | 2005 | 2025 | | `GEO` | 144 | 2009 | 2024 | | `KOR` | 144 | 2000 | 2025 | | `LKA` | 126 | 2010 | 2024 | | `TUR` | 126 | 2000 | 2013 | | `THA` | 126 | 2007 | 2024 | | `IND` | 120 | 1994 | 2025 | | ... | _15 more countries_ | | | ## Indicators (sample) - `EES_TEES_SEX_GEO_NB` — Employees by sex and rural / urban areas (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EES_TEES_SEX_GEO_NB` | | `indicator.label` | `string` | Indicator name in English | `Employees by sex and rural / urban ar…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `GEO_COV_NAT` | | `classif1.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `1709.649` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `B` | | `obs_status.label` | `string` | — | `Break in series` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (4 unique values): `SEX_T`, `SEX_M`, `SEX_F`, `SEX_O` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-ees-tees-sex-geo-nb-employees-by-sex-and-rural-urban-areas-thousands") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EES_TEES_SEX_GEO_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EES_TEES_SEX_GEO_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EES_TEES_SEX_GEO_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_ees_tees_sex_geo_nb_employees_by_sex_and_rural_urban_areas_thousands_2025, title = {Employees by sex and rural / urban areas (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_GEO_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-ees-tees-sex-geo-nb-employees-by-sex-and-rural-urban-areas-thousands}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_GEO_NB_

This dataset contains employee statistics for 30 Asian countries, disaggregated by sex and rural/urban areas, in thousands. It spans from 1970 to 2025 with 3,132 observations, focusing on the indicator EES_TEES_SEX_GEO_NB (Employees by sex and rural/urban areas). Sourced from the International Labour Organizations ILOSTAT database via API, the data is normalized and includes columns such as country codes, sex classifications (total, male, female, etc.), year, observed values, and data quality flags. It is designed for tabular classification, regression, and time-series forecasting tasks, providing a machine-learning-ready data layer for Asian labor market analysis.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-ees-tees-sex-geo-nb-employees-by-sex-and-rural-urban-areas-thousands 数据集图片
构建方式
数据集构建遵循严谨的统计标准化流程,其原始数据源自国际劳工组织(ILO)的中央统计数据库ILOSTAT,通过REST API接口直接提取指标EES_TEES_SEX_GEO_NB,并依据ISO3国家代码筛选出30个亚洲经济体,时间跨度为1970年至2025年。ILOSTAT采用国际劳工统计学家会议(ICLS)定义对各国劳动力调查、家庭收入调查等原始微观数据进行统一协调与校验,确保跨国可比性。Electric Sheep Asia在此基础上进行重新打包,将数据规范化为Parquet格式并发布至HuggingFace平台,每条观测均保留原始来源标签(source.label)和指标注释,实现了从原始调查到机器学习就绪数据集的透明转化。
特点
该数据集在劳动统计领域具有显著的时空覆盖优势与细分维度。它汇集了3,132条年度观测,涵盖30个亚洲国家,时间跨度超过半个世纪,提供了雇员人数按性别(总计、男性、女性)和城乡区域(国家层面)的双重分类。数据以千人为单位,包含ref_area、sex、classif1等丰富的分类变量以及观测状态标志(obs_status)和指标注释,能够支持对亚洲劳动力市场性别差异与城乡结构的长期演变分析。此外,数据质量注记明确区分了序列中断、临时性等状态,增强了研究结论的可靠性。
使用方法
研究人员可通过HuggingFace的datasets库便捷加载该数据集,使用load_dataset函数获取训练集并转换为Pandas DataFrame进行后续分析。典型操作包括按国家代码(如IDN)筛选特定经济体、按指标代码提取时间序列并可视化趋势,以及利用透视表功能将数据重塑为国家×年份矩阵以进行横向比较。该数据集适用于表格分类、回归及时间序列预测等任务,为劳动经济学、发展研究和区域政策评估提供了可直接使用的结构化数据基础。
背景与挑战
背景概述
国际劳工组织(ILO)自1919年成立以来,始终致力于全球劳动统计的标准化与传播,其ILOSTAT数据库是劳动统计领域的权威来源。该数据集由Electric Sheep Asia于2025年重新封装,涵盖1970至2025年间30个亚洲国家按性别与城乡划分的雇员数据,共计3,132条观测。数据集核心关注亚洲劳动力市场中的性别差异与城乡分化,为劳动经济学、社会政策及可持续发展目标(SDG)研究提供关键数据支撑,推动了区域劳动统计的透明化与可比性。
当前挑战
该数据集所应对的领域问题在于揭示亚洲地区雇员分布的性别与城乡异质性,这对理解劳动力市场不平等至关重要。构建过程中面临的挑战包括:跨国数据来源的异质性,需依赖ILO的协调标准以保障可比性;部分国家数据缺失或年份不连续,导致时间序列分析受限;城乡分类标准不一,可能引入测量偏差;此外,非正式就业与自雇人员常被排除,影响雇员统计的全面性。这些因素共同制约了数据在精细政策分析中的直接应用。
常用场景
经典使用场景
在劳动经济学与区域发展研究的交汇处,该数据集凭借其按性别及城乡地域双重维度拆分的雇员数量指标,成为刻画亚洲劳动力市场结构性特征的经典工具。研究者惯常利用其1970至2025年的长时序面板数据,构建城乡就业性别差距的演变轨迹,或借助表格分类与回归模型,检验城镇化进程中女性就业参与率的收敛与分化。其3,132条观测覆盖30个亚洲国家,为跨国比较与时间序列预测提供了均衡且标准化的分析基底。
衍生相关工作
围绕该数据集衍生的经典工作,多集中于劳动经济学与性别研究的交叉地带。一类研究以其为基准,检验亚洲城镇化进程中就业结构的性别极化假说;另一类则将其与工资、工时等ILOSTAT指标联结,构建多维劳动脆弱性指数。此外,部分机器学习研究利用其表格特征开展跨国就业分类与缺失值插补的方法探索,从而在数据工程与社会科学之间架设了方法论的桥梁。
数据集最近研究
最新研究方向
在劳动力市场性别平等与城乡发展差异的研究脉络中,该数据集为解析亚洲地区雇佣结构的时空演变提供了高粒度面板数据。近期前沿研究聚焦于利用其性别与城乡双重 disaggregation 维度,结合时间序列预测与因果推断方法,揭示结构性转型背景下女性就业的城乡收敛趋势、非正规经济冲击及政策干预效应。数据集覆盖三十国逾半世纪观测值,支撑了跨国比较与SDG体面劳动指标监测,对制定包容性就业政策、弥合区域发展鸿沟具有关键实证意义,亦推动了劳动经济学与空间计量交叉领域的范式创新。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务