遇见数据集

electricsheepeurope/europe-ilo-emp-pifl-sex-ocu-edu-rt-share-of-employment-outside-the-formal-sector-by-s

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - informal-economy - ilo - labour - employment pretty_name: "Share of employment outside the formal sector by sex, occupation and education (%) | Europe (ILOSTAT)" --- # Share of employment outside the formal sector by sex, occupation and education (%) | Europe (ILOSTAT) 🇪🇺 **84,958 observations** · **35 Europe countries** · **2003–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-84,958-blue) ![countries](https://img.shields.io/badge/countries-35-green) ![years](https://img.shields.io/badge/years-2003–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **84,958 observations** of `Informal economy` data across **35 Europe countries**, spanning **2003–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_PIFL_SEX_OCU_EDU_RT) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Informal economy ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EMP_PIFL_SEX_OCU_EDU_RT` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 35 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `RUS` | 6,971 | 2010 | 2025 | | `MDA` | 6,818 | 2003 | 2025 | | `BIH` | 5,942 | 2006 | 2024 | | `SRB` | 5,253 | 2007 | 2024 | | `MKD` | 4,916 | 2009 | 2025 | | `NLD` | 3,226 | 2007 | 2024 | | `ESP` | 2,679 | 2007 | 2024 | | `ITA` | 2,616 | 2007 | 2024 | | `FIN` | 2,562 | 2007 | 2024 | | `PRT` | 2,485 | 2007 | 2024 | | `GBR` | 2,440 | 2007 | 2018 | | `POL` | 2,414 | 2007 | 2024 | | `SVN` | 2,372 | 2007 | 2024 | | `SWE` | 2,229 | 2007 | 2024 | | `CZE` | 2,168 | 2007 | 2024 | | ... | _20 more countries_ | | | ## Indicators (sample) - `EMP_PIFL_SEX_OCU_EDU_RT` — Share of employment outside the formal sector by sex, occupation and education (%) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AUT` | | `ref_area.label` | `string` | Country name in English | `Austria` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:275` | | `source.label` | `string` | Source name in English | `HIES - EU Statistics on Income and Li…` | | `indicator` | `string` | ILOSTAT indicator code | `EMP_PIFL_SEX_OCU_EDU_RT` | | `indicator.label` | `string` | Indicator name in English | `Share of employment outside the forma…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `OCU_SKILL_TOTAL` | | `classif1.label` | `string` | — | `Occupation (Skill level): Total` | | `classif2` | `string` | Second classification variable where applicable | `EDU_AGGREGATE_TOTAL` | | `classif2.label` | `string` | — | `Education (Aggregate levels): Total` | | `time` | `int64` | Observation year | `2024` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `3.907` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `—` | | `note_classif.label` | `string` | — | `—` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513_T2:85` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-emp-pifl-sex-ocu-edu-rt-share-of-employment-outside-the-formal-sector-by-s") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EMP_PIFL_SEX_OCU_EDU_RT"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EMP_PIFL_SEX_OCU_EDU_RT") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EMP_PIFL_SEX_OCU_EDU_RT"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_emp_pifl_sex_ocu_edu_rt_share_of_employment_outside_the_formal_sector_by_s_2025, title = {Share of employment outside the formal sector by sex, occupation and education (%) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_PIFL_SEX_OCU_EDU_RT}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-emp-pifl-sex-ocu-edu-rt-share-of-employment-outside-the-formal-sector-by-s}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_PIFL_SEX_OCU_EDU_RT_

This dataset contains 84,958 observations from the International Labour Organization (ILO) ILOSTAT database, covering 35 European countries from 2003 to 2025. The core indicator is Share of employment outside the formal sector by sex, occupation and education (%) (indicator code: EMP_PIFL_SEX_OCU_EDU_RT). It focuses on the informal economy topic, providing employment share data disaggregated by dimensions such as sex (total, male, female), occupation, and education level. Data is sourced via the ILOSTAT API, filtered for European country codes, and harmonized, making it suitable for tasks like tabular classification, regression, and time-series forecasting.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-emp-pifl-sex-ocu-edu-rt-share-of-employment-outside-the-formal-sector-by-s 数据集图片
构建方式
该数据集依托国际劳工组织(ILO)的ILOSTAT中央统计数据库,通过REST API接口定向获取指标代码为EMP_PIFL_SEX_OCU_EDU_RT的原始记录,并依据ISO 3166-1 alpha-3国家代码筛选欧洲区域数据。ILO统计部门采用国际劳工统计学家会议(ICLS)标准定义对各国劳动力调查、家庭收入调查等微观数据进行统一调和,源数据来源在source.label列中予以标注,最终由Electric Sheep Europe完成Parquet格式的规范化打包与发布。
特点
数据集涵盖35个欧洲国家,时间跨度为2003年至2025年,共计84,958条观测记录。核心指标为非正规部门以外就业份额,按性别、职业与教育程度进行多维分解,sex列区分总计、男性与女性,classif1与classif2分别提供职业和教育的分类维度。数据以年度频率发布,包含来源代码、观测状态标志及系列断点注释等元数据字段,便于追溯与质量控制。
使用方法
研究者可通过HuggingFace的datasets库以load_dataset函数直接加载数据,并转换为Pandas数据框进行后续分析。典型操作包括按ref_area列筛选特定国家、按indicator列提取单一指标的时间序列、利用pivot_table生成国家与年份的矩阵视图,以及结合time列绘制趋势图。数据适用于表格分类、回归及时间序列预测等机器学习任务。
背景与挑战
背景概述
非正规经济部门的就业测度长期以来构成劳动统计领域的核心议题,其数据质量直接关系到体面劳动监测与可持续发展目标的评估成效。国际劳工组织(ILO)依托其中央统计数据库ILOSTAT,系统汇编了全球两百余个经济体的劳动力调查与家庭收入调查数据,并依据国际劳工统计学家会议(ICLS)的定义框架进行跨国协调。在此背景下,Electric Sheep Europe于2025年对ILOSTAT非正规部门就业指标进行重包装,构建了覆盖35个欧洲国家、时间跨度为2003至2025年、包含84,958条观测的专项数据集,按性别、职业与教育维度细分,为劳动经济学与非正规经济研究提供了标准化的面板数据基础。
当前挑战
非正规就业统计面临界定标准不统一与测量口径差异的固有困境,不同国家劳动力调查对非正规部门的操作化定义存在异质性,致使跨国比较的效度受到削弱。该数据集构建过程中的主要挑战包括:ILOSTAT多源数据整合需甄别各国调查在抽样设计、问卷措辞与参考期上的差异,并以ICLS框架实现事后协调;部分国家时间序列存在方法学修订所致断点,且在特定年份与细分维度上存在数值缺失或标记为不可靠的情形;性别、职业与教育三维交叉分类在部分国家未能完整发布,导致面板数据结构不平衡,对时间序列建模与跨国回归分析构成方法论挑战。
常用场景
经典使用场景
在劳动经济学与非正规经济研究领域,该数据集堪称一座结构化的实证宝库,其经典使用场景主要聚焦于对欧洲各国非正规部门就业份额的跨时空比较分析。研究者通常依据性别、职业与教育程度三重维度,提取2003至2025年间的连续观测值,构建面板数据模型或时间序列预测模型,以刻画不同国家非正规就业的演变轨迹。借助表格分类与回归任务,可深入探究性别差异、技能层级与教育水平如何交叠影响劳动者进入非正规部门的概率,进而为理解欧洲劳动力市场的结构性分化提供关键量化依据。
实际应用
在政策实践层面,该数据集为国际组织、欧洲各国劳工部门及社会保障机构提供了不可或缺的决策支持工具。依托其细分至性别、职业与教育维度的观测值,政策制定者可精准识别非正规就业的高发群体与高发区域,从而设计更具靶向性的社会保障扩面政策、技能培训计划与反贫困干预措施。同时,金融机构与智库亦可利用其时间序列特征,评估非正规经济对税收基数、社会保险可持续性及宏观经济增长的潜在影响,助力实现体面劳动与包容性增长的政策目标。
衍生相关工作
围绕该数据集,学术界与数据科学社区已衍生出一系列经典研究工作。在计量经济学领域,研究者利用其构建多层线性模型与工具变量回归,探讨非正规就业对收入不平等与性别工资差距的因果效应;在机器学习领域,学者将其作为表格分类与回归的基准数据,评估梯度提升、随机森林及深度神经网络在预测非正规就业份额上的表现。此外,部分研究将其与欧洲社会调查、收入与生活条件统计等微观数据链接,展开跨数据源的三角互证,推动了非正规经济测量方法论的精进与创新。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务