遇见数据集

electricsheepasia/asia-ilo-emp-pifl-sex-ocu-ins-rt-share-of-employment-outside-the-formal-sector-by-s

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - asia - ilostat - informal-economy - ilo - labour - employment pretty_name: "Share of employment outside the formal sector by sex, occupation and public/private sector | Asia (ILOSTAT)" --- # Share of employment outside the formal sector by sex, occupation and public/private sector | Asia (ILOSTAT) 🌏 **15,159 observations** · **25 Asia countries** · **2006–2025** · *Repackaged by [Electric Sheep Asia](https://huggingface.co/electricsheepasia)* ![rows](https://img.shields.io/badge/rows-15,159-blue) ![countries](https://img.shields.io/badge/countries-25-green) ![years](https://img.shields.io/badge/years-2006–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **15,159 observations** of `Informal economy` data across **25 Asia countries**, spanning **2006–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_PIFL_SEX_OCU_INS_RT) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Informal economy ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EMP_PIFL_SEX_OCU_INS_RT` and filtered to Asia ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 25 Asia countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `MNG` | 1,564 | 2006 | 2024 | | `LKA` | 1,398 | 2010 | 2024 | | `VNM` | 1,362 | 2007 | 2024 | | `TUR` | 1,344 | 2009 | 2024 | | `PSE` | 1,284 | 2010 | 2025 | | `PAK` | 1,120 | 2006 | 2025 | | `KGZ` | 1,094 | 2012 | 2023 | | `THA` | 998 | 2014 | 2024 | | `BGD` | 733 | 2010 | 2024 | | `BRN` | 664 | 2014 | 2024 | | `JOR` | 644 | 2017 | 2024 | | `GEO` | 460 | 2019 | 2024 | | `MMR` | 436 | 2015 | 2020 | | `ARM` | 404 | 2008 | 2017 | | `IDN` | 294 | 2016 | 2023 | | ... | _10 more countries_ | | | ## Indicators (sample) - `EMP_PIFL_SEX_OCU_INS_RT` — Share of employment outside the formal sector by sex, occupation and public/private sector (%) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AFG` | | `ref_area.label` | `string` | Country name in English | `Afghanistan` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:15715` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EMP_PIFL_SEX_OCU_INS_RT` | | `indicator.label` | `string` | Indicator name in English | `Share of employment outside the forma…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `OCU_SKILL_TOTAL` | | `classif1.label` | `string` | — | `Occupation (Skill level): Total` | | `classif2` | `string` | Second classification variable where applicable | `INS_SECTOR_TOTAL` | | `classif2.label` | `string` | — | `Institutional sector: Total` | | `time` | `int64` | Observation year | `2021` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `77.474` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `—` | | `note_classif.label` | `string` | — | `—` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_S3:8` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (4 unique values): `SEX_T`, `SEX_M`, `SEX_F`, `SEX_O` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepasia/asia-ilo-emp-pifl-sex-ocu-ins-rt-share-of-employment-outside-the-formal-sector-by-s") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python indonesia = df[df["ref_area"] == "IDN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EMP_PIFL_SEX_OCU_INS_RT"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EMP_PIFL_SEX_OCU_INS_RT") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EMP_PIFL_SEX_OCU_INS_RT"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{asia_ilo_emp_pifl_sex_ocu_ins_rt_share_of_employment_outside_the_formal_sector_by_s_2025, title = {Share of employment outside the formal sector by sex, occupation and public/private sector | Asia (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_PIFL_SEX_OCU_INS_RT}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Asia}, howpublished = {\url{https://huggingface.co/datasets/electricsheepasia/asia-ilo-emp-pifl-sex-ocu-ins-rt-share-of-employment-outside-the-formal-sector-by-s}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Asia repackaging. ## About Electric Sheep Electric Sheep Asia is part of the Electric Sheep mission: a unified, ML-ready data layer for Asia on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepasia](https://huggingface.co/electricsheepasia) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_PIFL_SEX_OCU_INS_RT_

This dataset contains 15,159 observations of informal economy data across 25 Asia countries, spanning from 2006 to 2025, covering one distinct indicator: the share of employment outside the formal sector by sex, occupation and public/private sector (EMP_PIFL_SEX_OCU_INS_RT). The data is sourced from the International Labour Organization (ILO) ILOSTAT database, pulled directly via REST API and filtered to Asia ISO3 country codes, with harmonization based on ICLS (International Conference of Labour Statisticians) definitions. It is structured in tabular format with columns such as country code, indicator code, sex disaggregation, time year, observed value, and includes source labels, quality flags, and notes. The dataset is repackaged by Electric Sheep Asia for machine learning readiness and is licensed under CC-BY-4.0, suitable for tasks like tabular classification, regression, and time-series forecasting.

提供机构:
electricsheepasia
搜集汇总
数据集介绍
electricsheepasia/asia-ilo-emp-pifl-sex-ocu-ins-rt-share-of-employment-outside-the-formal-sector-by-s 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT中央统计数据库,通过其REST API接口(https://rplumber.ilo.org/data/indicator?id=EMP_PIFL_SEX_OCU_INS_RT)直接拉取原始数据,并依据ISO3国家代码筛选出亚洲地区。ILOSTAT对各国劳动力调查、家庭收入调查等原始微观数据进行统一协调,采用国际劳工统计学家会议(ICLS)定义进行标准化处理,确保跨国可比性。数据经Electric Sheep Asia重新包装为HuggingFace数据集格式,保留来源标签以支持溯源,最终形成涵盖25个亚洲国家、2006至2025年间15,159条观测值的结构化数据。
特点
该数据集聚焦于非正规部门就业份额这一核心劳动统计指标,具有多维 disaggregation 特征:按性别(总计、男性、女性、其他)、职业技能水平以及公共/私营部门分类进行细分,全面刻画亚洲地区非正规经济的性别与职业结构差异。数据以年度频率呈现,覆盖25个亚洲国家,时间跨度近二十年,且包含观测状态标志(如 provisional、unreliable)和来源注释,便于评估数据质量。作为ILOSTAT的权威衍生数据集,其指标定义严格遵循国际标准,附带详尽的schema和溯源信息,适用于表格分类、回归及时间序列预测等任务。
使用方法
研究人员可通过HuggingFace datasets库以一行代码加载该数据集,并转换为Pandas DataFrame进行灵活分析。典型用法包括:利用`ref_area`字段筛选特定国家(如印度尼西亚),按`indicator`和`time`排序绘制单指标时间序列趋势图,或通过透视表将数据重塑为国家×年份矩阵以比较区域差异。数据集支持表格分类、回归和时间序列预测任务,用户还可结合`sex`、`classif1`等列进行分组统计分析。使用时需注意数据年度频率及来源选择逻辑,并遵循cc-by-4.0许可协议,同时引用ILO原始出处与Electric Sheep Asia的再包装工作。
背景与挑战
背景概述
非正规经济就业的规模与结构长期构成发展经济学与劳动经济学关注的核心议题,尤其在南亚、东南亚等非正规部门吸纳大量劳动力的区域,其测算与跨国可比性直接关系到体面劳动议程的推进。国际劳工组织(ILO)依托ILOSTAT数据库,依据国际劳工统计学家会议(ICLS)标准对各国劳动力调查等微观数据进行系统性协调,构建了按性别、职业及公共/私营部门分列的非正规部门外就业占比指标(EMP_PIFL_SEX_OCU_INS_RT)。Electric Sheep Asia于2025年前后将其重新打包发布,覆盖25个亚洲国家、2006—2025年共15159条观测,为劳动市场结构转型、性别就业差距及非正规性演变的研究提供了长时段、可复现的跨国面板数据。
当前挑战
该数据集所回应的领域难题在于非正规就业概念本身的多维性与测量标准随ICLS决议演进而变动,致使各国观测值在时间上难以直接比较,且部分国家仅零星发布特定性别或职业分组数据,形成大量结构性缺失。构建过程中的突出挑战包括:不同来源调查在抽样设计、覆盖范围与界定口径上差异显著,ILO虽选取最优来源并标注源标签,仍难以完全消除跨国异质性;部分观测被标记为不可靠或临时性,提示抽样误差与回应偏差的潜在影响;职业与机构部门分类维度并非在所有国家均予发布,导致分组序列非随机缺失,对分类与回归建模中的偏差校正与缺失机制假设提出更高要求;序列中记录的方法论修订断点亦要求时序分析中引入结构性变化识别与调整策略。
常用场景
经典使用场景
在劳动经济学与非正规经济研究领域,该数据集最经典的使用场景在于刻画亚洲地区非正规部门就业的性别分化与职业结构特征。研究者借助其涵盖25个亚洲国家、2006至2025年间15,159条观测记录,按性别、职业技能层级及公共/私营部门维度进行交叉分类,从而系统辨识非正规就业在不同人口群体中的分布态势。依托时间序列与面板数据方法,该数据集被广泛用于追踪各国非正规就业比重的演变轨迹,并支撑跨国比较与趋势检验。
解决学术问题
该数据集有效回应了非正规就业测度中长期存在的标准化与可比性难题。通过ILO统一采用国际劳工统计学家会议定义对原始调查微数据进行协调,它为学术研究提供了跨国家、跨年度可比的非正规就业指标,缓解了以往因各国调查口径差异而导致结论难以横向比较的困境。其意义在于将非正规就业从零星国别案例提升为区域面板分析对象,为检验制度变迁、经济周期与劳动力市场分割理论提供了实证基础,亦对可持续发展目标中体面劳动监测具有支撑价值。
衍生相关工作
围绕该数据集,衍生出一系列与非正规经济测度、性别就业差距及亚洲劳动力市场转型相关的研究工作。部分学者将其与ILOSTAT其他指标数据集联结,构建多维度体面劳动指数;亦有研究以其为基准,比较不同非正规就业定义对区域趋势估计的敏感性。这些工作推动了非正规就业统计方法的反思与改进,并促成针对亚洲新兴经济体劳动力市场二元结构的持续学术对话。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务