遇见数据集

electricsheepeurope/europe-ilo-emp-pifl-sex-est-mts-nb-employment-outside-the-formal-sector-by-sex-establ

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - informal-economy - ilo - labour - employment pretty_name: "Employment outside the formal sector by sex, establishment size and marital status (thousa | Europe (ILOSTAT)" --- # Employment outside the formal sector by sex, establishment size and marital status (thousa | Europe (ILOSTAT) 🇪🇺 **15,850 observations** · **4 Europe countries** · **2003–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-15,850-blue) ![countries](https://img.shields.io/badge/countries-4-green) ![years](https://img.shields.io/badge/years-2003–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **15,850 observations** of `Informal economy` data across **4 Europe countries**, spanning **2003–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_PIFL_SEX_EST_MTS_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Informal economy ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EMP_PIFL_SEX_EST_MTS_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 4 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `MDA` | 5,129 | 2003 | 2025 | | `MKD` | 4,182 | 2009 | 2025 | | `BIH` | 3,622 | 2006 | 2020 | | `SRB` | 2,917 | 2007 | 2020 | ## Indicators (sample) - `EMP_PIFL_SEX_EST_MTS_NB` — Employment outside the formal sector by sex, establishment size and marital status (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `BIH` | | `ref_area.label` | `string` | Country name in English | `Bosnia and Herzegovina` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:493` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EMP_PIFL_SEX_EST_MTS_NB` | | `indicator.label` | `string` | Indicator name in English | `Employment outside the formal sector …` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EST_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Establishment size (Aggregate): Total` | | `classif2` | `string` | Second classification variable where applicable | `MTS_AGGREGATE_TOTAL` | | `classif2.label` | `string` | — | `Marital status (Aggregate): Total` | | `time` | `int64` | Observation year | `2020` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `116.214` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `B` | | `obs_status.label` | `string` | — | `Break in series` | | `note_indicator` | `string` | — | `I20:4077_I11:264` | | `note_indicator.label` | `string` | — | `Employment definition: Excluding own-…` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-emp-pifl-sex-est-mts-nb-employment-outside-the-formal-sector-by-sex-establ") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EMP_PIFL_SEX_EST_MTS_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EMP_PIFL_SEX_EST_MTS_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EMP_PIFL_SEX_EST_MTS_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_emp_pifl_sex_est_mts_nb_employment_outside_the_formal_sector_by_sex_establ_2025, title = {Employment outside the formal sector by sex, establishment size and marital status (thousa | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_PIFL_SEX_EST_MTS_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-emp-pifl-sex-est-mts-nb-employment-outside-the-formal-sector-by-sex-establ}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_PIFL_SEX_EST_MTS_NB_

This dataset contains 15,850 observations of informal economy data across 4 Europe countries, spanning 2003–2025, covering 1 distinct indicator: EMP_PIFL_SEX_EST_MTS_NB — Employment outside the formal sector by sex, establishment size and marital status (thousands). The data is sourced from the International Labour Organization (ILO) ILOSTAT database, pulled via API and filtered to Europe ISO3 country codes. It includes a detailed schema with columns such as country code, indicator, sex disaggregation, time, and observed values, and is designed for tabular classification, regression, and time-series forecasting tasks. The data is annual frequency, with quality caveats like the use of ILO-selected best source. It supports quick loading and analysis for research on informal employment trends in Europe.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-emp-pifl-sex-est-mts-nb-employment-outside-the-formal-sector-by-sex-establ 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT中央统计数据库,通过其REST API接口直接提取指标EMP_PIFL_SEX_EST_MTS_NB的原始记录,并依据ISO3国家代码筛选出欧洲地区数据。ILOSTAT采用国际劳工统计学家会议(ICLS)标准对各国劳动力调查、家庭收入调查及行政记录等微观数据进行统一协调,数据经Electric Sheep Europe重新打包为Parquet格式并发布至HuggingFace平台。
特点
数据集涵盖四个欧洲国家(摩尔多瓦、北马其顿、波黑、塞尔维亚)在2003至2025年间的15,850条观测记录,聚焦非正规部门就业指标,按性别、机构规模及婚姻状况提供多维分解。数据以年度频率呈现,包含源标签、观测状态标志及详细注释字段,确保溯源透明。性别维度覆盖总计、男性与女性三类,分类变量按需非空,便于灵活筛选。
使用方法
研究者可借助HuggingFace的datasets库以load_dataset()函数快速加载数据,并转换为Pandas数据框进行后续分析。典型操作包括按国家代码过滤子集、提取单一指标的时间序列并可视化,或透视生成国家与年份的交叉矩阵。该数据集适用于非正规经济领域的表格分类、回归及时序预测任务,使用时需遵循CC-BY-4.0许可并同时引用ILO原始来源与Electric Sheep Europe的再包装工作。
背景与挑战
背景概述
非正规部门就业的规模与结构长期构成劳动经济学和发展经济学关注的核心议题,其测度直接关系到体面劳动议程的推进与社会保护政策的制定。国际劳工组织(ILO)依托其成员国劳动力调查与家庭收入调查等微观数据,经国际劳工统计学家会议(ICLS)定义加以协调,建立ILOSTAT这一全球劳动统计中枢。该数据集由Electric Sheep Europe于2025年基于ILOSTAT的REST API重新封装发布,覆盖摩尔多瓦、北马其顿、波黑与塞尔维亚四国,时间跨度为2003至2025年,包含15,850条观测,按性别、机构规模与婚姻状况对非正规部门就业进行多维分解,为转型经济体非正规就业的性别差异与结构演变研究提供了标准化、机器可读的数据基础。
当前挑战
该数据集所回应的领域问题在于如何以跨国可比的方式刻画非正规部门就业的规模与构成,这一问题的难度源于非正规就业概念界定在不同统计体系间的显著分歧,以及机构规模与婚姻状况等分类维度在各国调查中的非标准化处理。构建过程中面临的核心挑战包括:原始调查微数据在ICLS定义下的协调与口径统一,同一国家-年份存在多源数据时“最佳来源”的甄别与选择,断点、暂定值及不可靠观测等质量标志的系统性标注,以及部分分类维度仅在特定指标下发布所致的结构性缺失。此外,少数国家时间序列覆盖不均衡,也对跨期与跨国的稳健比较构成制约。
常用场景
经典使用场景
在劳动经济学与非正规经济研究领域,该数据集堪称剖析欧洲非正规部门就业性别差异的经典数据基础。其最典型的使用场景在于构建以性别、企业规模与婚姻状况为分类维度的面板数据模型,用以刻画2003至2025年间摩尔多瓦、北马其顿、波黑与塞尔维亚四国非正规就业的时序演变与结构特征。研究者常借助该数据集执行多维度交叉分析,例如比较不同婚姻状况下男女劳动者在微型与小型企业中的就业分布,或利用时间序列预测方法推演非正规就业的周期性波动。其15,850条观测与清晰的分类标签体系,令其成为劳动市场分层与脆弱就业群体识别研究的首选数据源之一。
衍生相关工作
围绕该数据集已衍生出一系列聚焦于非正规就业性别异质性的经典工作。部分研究以其为基准数据,结合欧洲社会调查微观数据构建多层次模型,验证家庭内部分工对女性非正规就业概率的调节效应;另有学者利用其企业规模维度,发展了非正规部门生产率估算的分解框架,将企业规模溢价与性别工资差距纳入统一分析。在方法论层面,该数据集常被用作时间序列预测与面板因果推断的测试平台,推动了小样本跨国比较中合成控制法与贝叶斯层次模型的迭代应用。此外,Electric Sheep Europe的标准化封装亦促进了劳动统计数据的可互操作实践,为后续多源ILO指标整合研究提供了可复用的模式参照。
数据集最近研究
最新研究方向
在全球非正规经济就业规模持续引发政策关切的背景下,该数据集聚焦欧洲四国(摩尔多瓦、北马其顿、波黑、塞尔维亚)非正规部门就业的性别差异、企业规模与婚姻状况交叉维度,为劳动经济学与性别研究提供了稀缺的时序微观基础。近期研究前沿趋向于运用该数据集进行非正规就业的性别分层机制分析,结合转型经济体制度变迁背景,探讨婚姻状况如何调节女性在微型企业中的就业脆弱性。该数据集亦被用于验证国际劳工组织关于非正规经济统计标准(ICLS)在东南欧地区的适用性,并支撑SDG目标8中体面工作指标的监测。其多维度分类体系为机器学习驱动的就业预测模型提供了结构化特征,推动了劳动统计与数据科学的交叉融合。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务