遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-eco-edu-nb-unemployment-of-previously-employed-persons-by-sex

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 100K<n<1M tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment of previously employed persons by sex, former economic activity and education | Europe (ILOSTAT)" --- # Unemployment of previously employed persons by sex, former economic activity and education | Europe (ILOSTAT) 🇪🇺 **151,287 observations** · **41 Europe countries** · **1976–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-151,287-blue) ![countries](https://img.shields.io/badge/countries-41-green) ![years](https://img.shields.io/badge/years-1976–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **151,287 observations** of `Unemployment` data across **41 Europe countries**, spanning **1976–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_ECO_EDU_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_ECO_EDU_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 41 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GBR` | 6,700 | 1987 | 2025 | | `ESP` | 6,374 | 1976 | 2025 | | `ITA` | 6,192 | 1992 | 2024 | | `FRA` | 6,076 | 1993 | 2022 | | `GRC` | 5,820 | 1993 | 2020 | | `DEU` | 5,503 | 1993 | 2024 | | `IRL` | 5,270 | 1988 | 2024 | | `PRT` | 5,225 | 1993 | 2024 | | `DNK` | 5,087 | 1993 | 2024 | | `BEL` | 5,077 | 1993 | 2024 | | `ROU` | 4,657 | 1994 | 2024 | | `CHE` | 4,625 | 1991 | 2025 | | `AUT` | 4,623 | 1985 | 2025 | | `HUN` | 4,603 | 1992 | 2024 | | `NLD` | 4,449 | 1996 | 2024 | | ... | _26 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_ECO_EDU_NB` — Unemployment of previously employed persons by sex, former economic activity and education (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:480` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_ECO_EDU_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment of previously employed p…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `ECO_SECTOR_TOTAL` | | `classif1.label` | `string` | — | `Economic activity (Broad sector): Total` | | `classif2` | `string` | Second classification variable where applicable | `EDU_AGGREGATE_TOTAL` | | `classif2.label` | `string` | — | `Education (Aggregate levels): Total` | | `time` | `int64` | Observation year | `2024` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `24.601` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:2138` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-eco-edu-nb-unemployment-of-previously-employed-persons-by-sex") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_ECO_EDU_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_ECO_EDU_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_ECO_EDU_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_eco_edu_nb_unemployment_of_previously_employed_persons_by_sex_2025, title = {Unemployment of previously employed persons by sex, former economic activity and education | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_ECO_EDU_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-eco-edu-nb-unemployment-of-previously-employed-persons-by-sex}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_ECO_EDU_NB_

This dataset, titled Unemployment of previously employed persons by sex, former economic activity and education | Europe (ILOSTAT), contains 151,287 observations across 41 European countries from 1976 to 2025. The core indicator is UNE_TUNE_SEX_ECO_EDU_NB, representing unemployment of previously employed persons by sex, former economic activity, and education (in thousands). Data is sourced from the International Labour Organization (ILO) ILOSTAT database, extracted via its REST API and filtered to European countries. The dataset includes detailed columns such as country code (ref_area), country name (ref_area.label), data source (source and source.label), indicator code and label (indicator and indicator.label), sex disaggregation (sex and sex.label), economic activity and education classifications (classif1, classif1.label, classif2, classif2.label), year (time), observed value (obs_value), and data status and note columns. Data is disaggregated by sex (total, male, female), economic activity, and education dimensions, suitable for tabular classification, regression, and time-series forecasting tasks. It is annual frequency and includes data quality caveats, such as the use of ILO-selected best source. The dataset is published in Parquet format for machine learning research.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-eco-edu-nb-unemployment-of-previously-employed-persons-by-sex 数据集图片
构建方式
该数据集由Electric Sheep Europe团队从ILOSTAT官方REST API直接拉取原始数据,并依据国际劳工统计学家会议(ICLS)定义进行统一清洗与整合。数据源涵盖各国劳动力调查、住户收入调查、机构调查及行政记录,ILO统计部门负责对原始调查微观数据进行标准化处理。团队进一步将数据范围限定至欧洲41个国家的ISO3代码,确保地域聚焦性,最终形成包含151,287条观测记录的高质量数据集,时间跨度为1976年至2025年。
特点
数据集以单一核心指标"先前就业人员失业人数"为核心,按性别、前经济活动和教育水平三个维度进行精细分层。其独特之处在于提供了完整的分类元数据,包括性别(总计、男性、女性)、经济活动部门及教育程度的详细标签,同时保留了观测状态标记与注释信息,便于用户识别数据质量(如临时性数据或中断序列)。此外,数据集标注了数据来源的溯源信息,增强了统计结果的透明度和可追溯性。
使用方法
用户可通过HuggingFace的`datasets`库以一行代码加载数据集,并将其转换为Pandas DataFrame进行灵活操作。典型应用包括按国家筛选子集、绘制单个指标的时间序列趋势图,以及将数据透视构建为国家×年份的矩阵结构,便于进行面板数据分析或机器学习建模。该数据集同时支持表格分类、回归以及时间序列预测等多种下游任务,研究者和开发者可直接利用标准化接口开展欧洲劳动力市场的实证研究。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)统计数据库ILOSTAT于2025年整理发布,经Electric Sheep Europe重新打包并托管于HuggingFace平台。其核心研究问题聚焦于欧洲地区先前就业人员的失业状况,通过性别、先前经济活动及教育程度等多维度分类,揭示劳动力市场的结构性特征。数据集覆盖41个欧洲国家,时间跨度自1976年至2025年,包含151,287条观测记录,为劳动经济学、失业动态监测及跨国比较研究提供了宝贵的高分辨率时空数据。自发布以来,该数据集在劳动力市场分析、政策评估及机器学习模型训练等领域展现出重要影响力,尤其为探究欧洲各国失业模式的异质性及其与教育、产业结构的关联奠定了基础。
当前挑战
该数据集所面临的挑战首先源于劳动统计领域固有的复杂性:失业的定义与测量标准虽由国际劳工统计学家会议(ICLS)统一协调,但各国在数据收集方法、调查频率及分类体系上仍存在差异,导致跨国比较时需谨慎处理口径不一致的问题。构建过程中,数据整合面临多重困难:原始数据来自不同国家的劳动力调查、行政记录等多种来源,需通过ILO的“最佳来源”选择机制进行协调,但部分观测值被标记为“不可靠”或存在序列断裂,提示方法论修订对数据连续性的影响。此外,分类变量(如教育程度)的非标准定义及缺失值处理,进一步增加了数据清洗与归一化的技术挑战,要求用户在设计分析模型时充分考虑这些潜在偏差与局限性。
常用场景
经典使用场景
在欧洲劳动力市场研究领域,该数据集被广泛用于分析不同性别、先前经济活动部门及教育背景下的失业人员结构特征。研究者常将其作为面板数据,探究欧洲41国从1976年至2025年间失业现象的时空演变,尤其关注性别差异与教育水平如何交织影响失业风险。通过整合ILOSTAT标准化的劳动力调查数据,该数据集使得跨国、跨时期的比较分析成为可能,是劳动经济学、人口社会学等领域构建回归模型或时间序列预测的经典数据基石。
实际应用
在公共政策与商业决策领域,该数据集的实际应用价值显著。欧盟及各成员国劳动部门可基于其精细的分层数据,设计更具靶向性的再就业培训计划,例如识别受教育程度较低或特定前行业(如制造业)中女性失业者的集中趋势。国际劳工组织亦能利用这些年度观测值,监测区域间劳动权益保障目标的实现进度。对于跨国企业而言,依据不同国家、性别与教育组合的失业率变化,可优化在欧洲市场的人力资源布局与投资风险评估。
衍生相关工作
基于该数据,衍生出若干具有影响力的经典研究路径。其一,学者们常将其与宏观经济学指标耦合,构建欧洲国家失业率的动态面板模型,并引入性别与教育虚拟变量以检验结构性假说。其二,时间序列分解方法被广泛用于提取长周期失业趋势,并量化1990年代东欧转型及2008年金融危机对特定人群的差异化冲击。此外,该数据集与ILO其他就业质量数据(如工作贫困率、非正规就业比例)的联合分析,催生了一系列关于体面劳动(Decent Work)综合评估的计量框架,成为欧洲社会政策评估领域的重要参考工具。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务