遇见数据集

electricsheepasia/asia-ilo-ees-tees-sex-eco-geo-nb-employees-by-sex-economic-activity-and-rural-urban

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - asia - ilostat - employees - ilo - labour - employment pretty_name: "Employees by sex, economic activity and rural / urban areas (thousands) | Asia (ILOSTAT)" --- # Employees by sex, economic activity and rural / urban areas (thousands) | Asia (ILOSTAT) 🌏 **36,977 observations** · **30 Asia countries** · **1970–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-36,977-blue) ![countries](https://img.shields.io/badge/countries-30-green) ![years](https://img.shields.io/badge/years-1970–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **36,977 observations** of `Employees` data across **30 Asia countries**, spanning **1970–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_ECO_GEO_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Employees ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EES_TEES_SEX_ECO_GEO_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 30 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `PSE` | 3,229 | 2000 | 2022 | | `CYP` | 2,785 | 1999 | 2024 | | `IDN` | 2,680 | 2000 | 2023 | | `KHM` | 2,149 | 1996 | 2023 | | `MNG` | 2,061 | 2003 | 2024 | | `VNM` | 1,970 | 2007 | 2024 | | `PAK` | 1,866 | 2005 | 2025 | | `ARM` | 1,818 | 2007 | 2023 | | `KOR` | 1,728 | 2000 | 2025 | | `GEO` | 1,719 | 2009 | 2024 | | `THA` | 1,708 | 2007 | 2024 | | `TUR` | 1,544 | 2000 | 2013 | | `LKA` | 1,539 | 2010 | 2024 | | `IND` | 1,454 | 1994 | 2025 | | `PHL` | 1,332 | 2012 | 2023 | | ... | _15 more countries_ | | | ## Indicators (sample) - `EES_TEES_SEX_ECO_GEO_NB` — Employees by sex, economic activity and rural / urban areas (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EES_TEES_SEX_ECO_GEO_NB` | | `indicator.label` | `string` | Indicator name in English | `Employees by sex, economic activity a…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `ECO_SECTOR_TOTAL` | | `classif1.label` | `string` | — | `Economic activity (Broad sector): Total` | | `classif2` | `string` | Second classification variable where applicable | `GEO_COV_NAT` | | `classif2.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `1709.649` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `float64` | — | `—` | | `note_classif.label` | `float64` | — | `—` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (4 unique values): `SEX_T`, `SEX_M`, `SEX_F`, `SEX_O` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-ees-tees-sex-eco-geo-nb-employees-by-sex-economic-activity-and-rural-urban") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EES_TEES_SEX_ECO_GEO_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EES_TEES_SEX_ECO_GEO_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EES_TEES_SEX_ECO_GEO_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_ees_tees_sex_eco_geo_nb_employees_by_sex_economic_activity_and_rural_urban_2025, title = {Employees by sex, economic activity and rural / urban areas (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_ECO_GEO_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-ees-tees-sex-eco-geo-nb-employees-by-sex-economic-activity-and-rural-urban}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_ECO_GEO_NB_

This dataset contains employee data from the International Labour Organization (ILO) ILOSTAT database for 30 Asian countries spanning 1970 to 2025, with 36,977 observations. The core indicator is EES_TEES_SEX_ECO_GEO_NB, which represents employees by sex, economic activity, and rural/urban areas (in thousands). Data includes disaggregation dimensions such as sex (total, male, female, other), economic activity (by broad sector), and area type (e.g., national). Organized in tabular format, the dataset features columns for country codes, years, observed values, data sources, and quality flags, suitable for tabular classification, regression, and time-series forecasting tasks. Data is sourced via the ILOSTAT API, filtered for Asia, and normalized for consistency, providing a machine learning-ready layer for Asian labor statistics.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-ees-tees-sex-eco-geo-nb-employees-by-sex-economic-activity-and-rural-urban 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的核心统计数据库ILOSTAT,通过其REST API接口直接提取指标EES_TEES_SEX_ECO_GEO_NB的原始数据,并依据ISO3国家代码筛选出亚洲地区30个国家的记录。ILOSTAT采用国际劳工统计学家会议(ICLS)的定义对各国劳动力调查等微观数据进行标准化处理,数据集保留了来源标签以确保可追溯性,最终由Electric Sheep Asia重新打包为Parquet格式并发布。
特点
数据集涵盖1970年至2025年间亚洲30个国家的36,977条观测记录,专注于按性别、经济活动类别及城乡区域划分的雇员数量(千人)。数据以年度频率呈现,包含性别(总数、男性、女性等)、经济活动分类和城乡覆盖等多维度分解字段,同时提供观测状态标志和来源注释,便于评估数据质量与可靠性。
使用方法
研究者可通过HuggingFace的datasets库以一行代码加载数据集,并转换为Pandas数据框进行灵活分析。支持按国家、指标或时间筛选子集,亦可绘制时间序列曲线或构建国家×年份矩阵,适用于表格分类、回归及时间序列预测等机器学习任务,为亚洲劳动力市场研究提供便捷的数据支撑。
背景与挑战
背景概述
国际劳工组织(ILO)长期致力于全球劳动力市场统计监测,其ILOSTAT数据库为劳动经济学与就业结构研究提供了权威数据基础。该数据集由Electric Sheep Asia于2025年从ILOSTAT REST API抓取并重新封装,涵盖1970至2025年间30个亚洲国家的36,977条观测记录,聚焦按性别、经济活动及城乡区域划分的雇员规模。数据集旨在回应亚洲地区就业结构变迁中的性别差异与城乡分化等核心议题,为劳动经济学、区域发展研究及政策评估提供了长时序、多维度、可机读的标准化数据资源,对推动亚洲劳动力市场实证研究具有重要支撑价值。
当前挑战
该数据集所回应的领域问题在于如何精确刻画亚洲各国就业结构的性别与城乡异质性,其难点源于各国劳动力调查体系在统计口径、抽样方法及分类标准上的长期不统一,导致跨国比较面临可比性障碍。构建过程中,ILOSTAT虽依托国际劳工统计学家会议(ICLS)定义进行协调,但原始微数据经多源汇集与‘最佳来源’筛选,仍存在序列断裂、观测状态标记为不可靠及部分国家年份覆盖不完整等挑战。此外,性别与城乡分类维度仅在特定指标下非空,增加了数据建模与时间序列预测的复杂性。
常用场景
经典使用场景
在劳动经济学与性别研究的交叉领域,该数据集最为经典的使用场景在于刻画亚洲各国男女雇员在经济活动与城乡地域维度上的分布特征及其时序演变。研究者常以性别、经济活动部门与城乡属性为分组变量,构建面板数据模型或时间序列模型,用以揭示女性劳动参与率的结构性变动、城乡就业机会的非均衡分布以及经济结构转型对雇员构成的深层影响。凭借1970至2025年的长时段覆盖与30个亚洲国家的横截面广度,该数据集为跨国比较与长期趋势分析提供了不可多得的量化基础。
解决学术问题
该数据集有效回应了劳动统计学中长期存在的若干学术难题,包括性别就业差距的测度与分解、非正规就业与正规就业的城乡分化、以及经济周期波动对男女雇员数量的异质性冲击。借助标准化的ICLS定义与ILO统一协调方法,它缓解了跨国劳动统计口径不一致的痼疾,使研究者得以在可比框架下检验性别隔离假说、结构性转型理论及城乡二元劳动力市场模型。其意义在于将碎片化的国别调查数据整合为可复现的科研资源,为循证政策评估与劳动市场理论发展奠定了坚实的数据根基。
衍生相关工作
围绕该数据集及其母体ILOSTAT体系,学界与开源社区衍生出诸多经典工作。比较劳动经济学文献中,基于ILO数据的跨国性别就业差距分解研究、城乡劳动力市场分割的计量分析以及结构转型与女性就业关系的长时段考察均大量援引此类指标。在数据科学领域,Electric Sheep Asia的规范化重打包催生了亚洲劳动市场数据湖、自动化时间序列预测基准以及表格数据插补与因果推断方法评测等一系列衍生资源。这些工作共同拓展了ILO统计数据在机器学习与政策分析中的复用边界。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务