遇见数据集

electricsheepeurope/europe-ilo-ees-tees-sex-ifl-ocu-nb-employees-by-sex-informal-formal-job-and-occupatio

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - informal-economy - ilo - labour - employment pretty_name: "Employees by sex, informal/formal job and occupation (thousands) | Europe (ILOSTAT)" --- # Employees by sex, informal/formal job and occupation (thousands) | Europe (ILOSTAT) 🇪🇺 **12,314 observations** · **5 Europe countries** · **2003–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-12,314-blue) ![countries](https://img.shields.io/badge/countries-5-green) ![years](https://img.shields.io/badge/years-2003–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **12,314 observations** of `Informal economy` data across **5 Europe countries**, spanning **2003–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_IFL_OCU_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Informal economy ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EES_TEES_SEX_IFL_OCU_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 5 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `MDA` | 3,032 | 2003 | 2025 | | `BIH` | 2,548 | 2006 | 2024 | | `SRB` | 2,481 | 2007 | 2025 | | `MKD` | 2,237 | 2009 | 2025 | | `RUS` | 2,016 | 2010 | 2025 | ## Indicators (sample) - `EES_TEES_SEX_IFL_OCU_NB` — Employees by sex, informal/formal job and occupation (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `BIH` | | `ref_area.label` | `string` | Country name in English | `Bosnia and Herzegovina` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:493` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EES_TEES_SEX_IFL_OCU_NB` | | `indicator.label` | `string` | Indicator name in English | `Employees by sex, informal/formal job…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `IFL_NATURE_TOTAL` | | `classif1.label` | `string` | — | `Nature of job: Total` | | `classif2` | `string` | Second classification variable where applicable | `OCU_SKILL_TOTAL` | | `classif2.label` | `string` | — | `Occupation (Skill level): Total` | | `time` | `int64` | Observation year | `2024` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `1069.429` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `—` | | `note_classif.label` | `string` | — | `—` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-ees-tees-sex-ifl-ocu-nb-employees-by-sex-informal-formal-job-and-occupatio") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EES_TEES_SEX_IFL_OCU_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EES_TEES_SEX_IFL_OCU_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EES_TEES_SEX_IFL_OCU_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_ees_tees_sex_ifl_ocu_nb_employees_by_sex_informal_formal_job_and_occupatio_2025, title = {Employees by sex, informal/formal job and occupation (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_IFL_OCU_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-ees-tees-sex-ifl-ocu-nb-employees-by-sex-informal-formal-job-and-occupatio}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_IFL_OCU_NB_

This dataset contains labor statistics on the informal economy in Europe, specifically the indicator Employees by sex, informal/formal job and occupation (thousands). It includes 12,314 observations covering 5 European countries (MDA, BIH, SRB, MKD, RUS) from 2003 to 2025. The data is sourced from the International Labour Organizations (ILO) ILOSTAT database, retrieved via its REST API and filtered for European country codes. The dataset is in tabular format with columns such as country code, country name, data source, indicator code, indicator label, sex disaggregation, classification variables, observation year, observed value, observation status flags, etc. Data is provided at annual frequency and includes quality caveats, such as some observations being flagged as provisional or unreliable. It is suitable for tasks like tabular classification, tabular regression, and time-series forecasting, and can be used to analyze informal employment in European labor markets.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-ees-tees-sex-ifl-ocu-nb-employees-by-sex-informal-formal-job-and-occupatio 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的中央统计数据库ILOSTAT,通过其REST API接口直接抽取指标EES_TEES_SEX_IFL_OCU_NB的原始记录,并依据ISO 3166-1 alpha-3国家代码筛选出欧洲区域样本。原始调查微观数据经ILO统计局依据国际劳工统计学家会议(ICLS)定义进行标准化调和,来源信息以source.label字段标注以确保可追溯性,最终由Electric Sheep Europe重新封装为HuggingFace数据集格式。
特点
数据集涵盖5个欧洲国家(摩尔多瓦、波黑、塞尔维亚、北马其顿、俄罗斯),时间跨度为2003至2025年,共12,314条观测记录。核心指标按性别(总计、男性、女性)以及非正规/正规就业和职业分类进行多维 disaggregation,包含国家代码、来源标签、指标代码、分类变量、观测值及状态标志等丰富字段,支持表格分类、回归与时间序列预测等多种任务,以Parquet格式存储便于高效加载。
使用方法
研究者可通过HuggingFace的datasets库以load_dataset()函数直接加载数据集,并转换为Pandas DataFrame进行灵活操作。典型用法包括按ref_area字段筛选特定国家、依据indicator字段提取单一指标的时间序列、利用pivot_table生成国家×年份矩阵以辅助趋势分析,亦可结合obs_status标志评估数据可靠性,从而满足劳动经济学、非正规就业研究及机器学习建模等多场景需求。
背景与挑战
背景概述
非正规经济就业的测度长期构成劳动统计领域的核心难题,其规模与结构直接关乎社会保障覆盖、体面劳动监测及可持续发展目标的评估。国际劳工组织自成立以来持续推动劳动统计标准化,其ILOSTAT数据库汇集全球二百余个经济体的劳动力调查与行政记录,构成该领域最权威的数据基础设施。本数据集由Electric Sheep Europe于2025年前后重新封装发布,原始数据取自ILO的ILOSTAT平台,覆盖摩尔多瓦、波黑、塞尔维亚、北马其顿及俄罗斯五个欧洲国家,时间跨度为2003至2025年,共12,314条观测,按性别、非正规/正规就业性质及职业分类三个维度记录雇员人数。该数据集将分散的国别劳动统计整合为机器学习就绪的表格结构,为转型经济体非正规就业的量化研究提供了可复用的数据基础。
当前挑战
非正规就业统计的根本困难在于概念界定与测量口径的不统一,国际劳工统计学家会议虽已确立标准定义,但各国劳动力调查在抽样设计、问卷措辞及非正规部门识别方法上存在显著差异,致使跨国可比性受限。就本数据集而言,其构建面临多重具体挑战:一是覆盖范围狭隘,仅纳入五个欧洲国家,难以支撑区域层面的稳健推断;二是数据缺失与序列断裂频发,各国起始年份参差不齐,且存在方法论修订所致的断点,如指标注释中标记的序列断裂;三是观测状态标注显示部分数值被评定为不可靠,影响建模质量;四是分类维度仅在特定指标下发布,造成不同维度组合间的样本量不均衡,对细粒度分析构成制约。
常用场景
经典使用场景
在劳动力市场统计与就业结构分析领域,该数据集最为经典的使用场景在于刻画欧洲非正规经济中按性别与职业划分的雇员分布格局。研究者可借助其跨越2003至2025年的年度观测值,对摩尔多瓦、波黑、塞尔维亚、北马其顿及俄罗斯等国的非正规与正规就业构成进行纵向追踪与横向比较。通过对职业分类维度的细分,能够揭示不同技能层级中非正规就业的性别差异,进而为研判转型经济体劳动力市场的结构性特征提供实证基础。
实际应用
在政策制定与社会保障领域,该数据集为识别非正规就业脆弱群体提供了量化依据。劳动监察部门可依据按性别和职业细分的非正规雇员规模,评估现行劳动法规的覆盖盲区,并针对性地设计扩展社会保障覆盖面的干预措施。国际组织与智库亦可利用其时间序列特性,监测非正规经济随经济周期与制度变迁的演变轨迹,为体面劳动议程的国别进展评估提供数据支撑。
衍生相关工作
围绕该数据集及其所依托的国际劳工组织统计体系,衍生出若干具有影响力的研究方向。其中包括对非正规就业与贫困脆弱性关联的跨国计量分析、结合劳动力调查微观数据展开的方法论校准研究,以及基于面板数据探讨制度质量对非正规经济规模影响的比较政治经济学文献。这些工作在不同程度上援引该数据集的结构化指标,拓展了非正规就业研究的理论深度与政策相关性。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务