遇见数据集

electricsheepasia/asia-ilo-ees-tees-sex-oc2-nb-employees-by-sex-and-occupation-isco-level-2-thous

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - asia - ilostat - employees - ilo - labour - employment pretty_name: "Employees by sex and occupation - ISCO level 2 (thousands) | Asia (ILOSTAT)" --- # Employees by sex and occupation - ISCO level 2 (thousands) | Asia (ILOSTAT) 🌏 **36,717 observations** · **35 Asia countries** · **1996–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-36,717-blue) ![countries](https://img.shields.io/badge/countries-35-green) ![years](https://img.shields.io/badge/years-1996–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **36,717 observations** of `Employees` data across **35 Asia countries**, spanning **1996–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_OC2_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Employees ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EES_TEES_SEX_OC2_NB` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 35 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `CYP` | 2,594 | 1999 | 2024 | | `PSE` | 2,372 | 2000 | 2025 | | `TUR` | 2,312 | 2000 | 2024 | | `MNG` | 2,159 | 2000 | 2024 | | `VNM` | 2,145 | 2007 | 2024 | | `KHM` | 2,010 | 1996 | 2023 | | `IRN` | 1,839 | 2006 | 2024 | | `THA` | 1,763 | 2000 | 2024 | | `PHL` | 1,678 | 2007 | 2023 | | `LKA` | 1,673 | 2010 | 2024 | | `ISR` | 1,586 | 2012 | 2024 | | `PAK` | 1,578 | 2006 | 2025 | | `GEO` | 1,540 | 2009 | 2024 | | `JPN` | 1,392 | 2000 | 2023 | | `KGZ` | 1,390 | 2010 | 2023 | | ... | _20 more countries_ | | | ## Indicators (sample) - `EES_TEES_SEX_OC2_NB` — Employees by sex and occupation - ISCO level 2 (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EES_TEES_SEX_OC2_NB` | | `indicator.label` | `string` | Indicator name in English | `Employees by sex and occupation - ISC…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `OC2_ISCO08_TOTAL` | | `classif1.label` | `string` | — | `Occupation (ISCO-08), 2 digit level: …` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `1709.649` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C4:6305` | | `note_classif.label` | `string` | — | `Nonstandard occupation: Including ass…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (4 unique values): `SEX_T`, `SEX_M`, `SEX_F`, `SEX_O` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-ees-tees-sex-oc2-nb-employees-by-sex-and-occupation-isco-level-2-thous") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EES_TEES_SEX_OC2_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EES_TEES_SEX_OC2_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EES_TEES_SEX_OC2_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_ees_tees_sex_oc2_nb_employees_by_sex_and_occupation_isco_level_2_thous_2025, title = {Employees by sex and occupation - ISCO level 2 (thousands) | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_OC2_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-ees-tees-sex-oc2-nb-employees-by-sex-and-occupation-isco-level-2-thous}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_OC2_NB_

This dataset contains 36,717 observations of employee data across 35 Asian countries, spanning from 1996 to 2025, covering 1 distinct indicator: Employees by sex and occupation - ISCO level 2 (thousands). The data is sourced from the International Labour Organization (ILO) ILOSTAT database, retrieved via API and filtered to Asian countries. It is organized in tabular format with columns such as country code, country name, data source, indicator code, indicator label, sex classification, time year, observed value, etc., supporting disaggregation by sex (total, male, female, other) and occupation classification. The data is annual frequency and suitable for tasks like tabular classification, regression, and time-series forecasting. The dataset is repackaged by Electric Sheep Asia in Parquet format for machine learning readiness.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-ees-tees-sex-oc2-nb-employees-by-sex-and-occupation-isco-level-2-thous 数据集图片
构建方式
该数据集源自国际劳工组织ILOSTAT统计数据库,由Electric Sheep Asia通过其数据管道从ILOSTAT REST API接口直接拉取原始数据,并依据亚洲ISO3国家代码进行筛选与重组。ILOSTAT汇聚各国劳动力调查、家庭收入调查及行政记录等多元数据源,依据国际劳工统计学家会议定义对原始微观数据进行标准化调和,确保指标口径一致。该数据集在保留源标签以追溯数据来源的同时,将原始数据打包为Parquet格式并统一模式,最终形成覆盖35个亚洲国家、时间跨度为1996至2025年的结构化表格数据。
特点
数据集包含36,717条观测记录,涵盖35个亚洲国家,时间序列跨越近三十年。其核心指标为按性别与职业分类的雇员人数,职业分类采用国际标准职业分类ISCO二级编码,性别维度细分为总计、男性、女性及其他四类。数据以年度频率发布,每个国家与年份组合采用ILO甄选的最优数据源。表格中设有来源标签、观测状态标志及多重注释字段,用于标记数据质量与口径变动,便于用户辨识数据可靠性与方法学调整。
使用方法
用户可通过HuggingFace的datasets库调用load_dataset函数加载该数据集,并利用to_pandas方法转换为数据框进行后续分析。典型用法包括按ref_area列筛选特定国家,或按indicator列提取单一指标的时间序列并绘制趋势图。借助pivot_table函数可将数据重塑为国家与年份的交叉矩阵,便于开展跨国比较或面板数据分析。数据以CC-BY-4.0协议开放,使用时须同时引用国际劳工组织原始来源与Electric Sheep Asia的再包装工作。
背景与挑战
背景概述
伴随全球化进程的深入,劳动力市场性别结构与职业分布的比较研究日益成为劳动经济学与发展经济学的核心议题。国际劳工组织(ILO)长期致力于构建跨国可比的劳动力统计数据体系,其ILOSTAT数据库整合了各国劳动力调查与住户调查等多源微观数据,并依据国际劳工统计学家会议(ICLS)定义进行标准化处理。本数据集由Electric Sheep Asia于2025年重新打包发布,源自ILOSTAT官方API,涵盖35个亚洲国家、1996至2025年间按性别与ISCO-08二位职业分类统计的雇员人数,共36,717条观测,为亚洲地区性别职业隔离、就业结构变迁等研究提供了坚实的数据基础。
当前挑战
该数据集所回应的领域问题在于如何在跨国、跨时间的异质环境中实现职业与性别维度上雇员统计的可比性,这一目标面临诸多固有挑战。首要难题在于各国劳动力调查在抽样设计、职业编码实践及统计口径上存在显著差异,尽管ILO进行了统一调和,但系列中断、方法修订与非标准职业归类等情形仍普遍存在,观测状态标记亦提示部分数据的可靠性有限。其次,构建过程中需应对数据稀疏性、缺失值以及不同国家年份覆盖不均衡等问题,且性别维度中‘其他’类别样本极为有限,为分类与预测建模带来额外困难,对模型的泛化能力提出了更高要求。
常用场景
经典使用场景
在劳动经济学与性别研究的交叉领域,该数据集最为经典的运用当属面向亚洲地区雇佣结构的纵贯比较分析。研究者依托国际劳工组织标准化后的ISCO二级职业分类框架,辅以性别维度(总计、男性、女性及其他)的细分标识,得以在1996至2025年的三十载时间跨度内,精准刻画35个亚洲经济体的受雇人口在各类职业中的规模分布与演变轨迹。通过构造国家—年份—职业—性别的四维面板结构,该数据集支撑了针对职业隔离程度、女性在专业与管理岗位中占比变化以及非标准职业扩张等议题的系统性实证考查。
实际应用
在实务层面,该数据集为国际组织、政府劳动部门及智库机构提供了可资依托的证据基础。其具体的应用场景包括:监测亚洲各国在可持续发展目标(SDGs)第八项“体面工作”框架下性别平等就业指标的推进状况,识别女性在文书支持、服务销售及手工艺等职业类别中的过度集中现象,并为职业培训政策与反歧视立法的设计提供量化参照。同时,凭借Parquet格式的机器学习友好封装,该数据集亦便于数据科学从业者快速构建分类、回归或时间序列预测模型,用以评估不同国家未来就业结构变化的可能路径。
衍生相关工作
该数据集衍生的相关研究工作主要沿着两条脉络展开。其一是在ILOSTAT原始指标基础上,结合其他劳动力市场数据集(如失业率、工资水平、非正规就业率)进行的整合性分析,试图揭示受雇结构与劳动报酬之间的联动机制;其二是区域层面比较研究的深化,部分学者利用该数据集中亚洲各国的异质性开展分组回归,探讨经济发展水平、女性劳动参与率与职业隔离指数之间的关联。此外,Electric Sheep Asia的标准化重封装实践本身也构成了一项值得关注的工程性贡献,其为后续其他地区或指标的平行数据集构建提供了可复制的方法论范式。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务