遇见数据集

electricsheepeurope/europe-ilo-une-deap-sex-edu-geo-rt-unemployment-rate-by-sex-education-and-rural-urban

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment rate by sex, education and rural / urban areas (%) | Europe (ILOSTAT)" --- # Unemployment rate by sex, education and rural / urban areas (%) | Europe (ILOSTAT) 🇪🇺 **76,163 observations** · **39 Europe countries** · **1987–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-76,163-blue) ![countries](https://img.shields.io/badge/countries-39-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **76,163 observations** of `Unemployment` data across **39 Europe countries**, spanning **1987–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_EDU_GEO_RT) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_DEAP_SEX_EDU_GEO_RT` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 39 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GRC` | 4,805 | 1987 | 2025 | | `FRA` | 3,128 | 1993 | 2024 | | `ITA` | 2,895 | 1992 | 2024 | | `IRL` | 2,886 | 1993 | 2024 | | `SWE` | 2,776 | 1995 | 2024 | | `PRT` | 2,774 | 1992 | 2024 | | `ESP` | 2,748 | 1992 | 2024 | | `DEU` | 2,677 | 1992 | 2024 | | `BEL` | 2,647 | 1992 | 2024 | | `DNK` | 2,611 | 1992 | 2024 | | `NLD` | 2,592 | 1996 | 2024 | | `AUT` | 2,574 | 1995 | 2025 | | `SRB` | 2,370 | 2007 | 2025 | | `HUN` | 2,236 | 2001 | 2024 | | `FIN` | 2,131 | 1995 | 2024 | | ... | _24 more countries_ | | | ## Indicators (sample) - `UNE_DEAP_SEX_EDU_GEO_RT` — Unemployment rate by sex, education and rural / urban areas (%) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:7401` | | `source.label` | `string` | Source name in English | `HIES - Living Standards Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_DEAP_SEX_EDU_GEO_RT` | | `indicator.label` | `string` | Indicator name in English | `Unemployment rate by sex, education a…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `classif2` | `string` | Second classification variable where applicable | `GEO_COV_NAT` | | `classif2.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2012` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `22.032` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:5578` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-deap-sex-edu-geo-rt-unemployment-rate-by-sex-education-and-rural-urban") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_DEAP_SEX_EDU_GEO_RT"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_DEAP_SEX_EDU_GEO_RT") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_DEAP_SEX_EDU_GEO_RT"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_deap_sex_edu_geo_rt_unemployment_rate_by_sex_education_and_rural_urban_2025, title = {Unemployment rate by sex, education and rural / urban areas (%) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_EDU_GEO_RT}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-deap-sex-edu-geo-rt-unemployment-rate-by-sex-education-and-rural-urban}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_EDU_GEO_RT_

This dataset contains unemployment rate data from the International Labour Organization (ILO) ILOSTAT database, specifically for Europe. It covers 39 European countries from 1987 to 2025, with 76,163 observations. The key indicator is Unemployment rate by sex, education and rural / urban areas (%), disaggregated by sex (total, male, female), education (aggregate levels), and area type (e.g., national). Data is sourced from national labour force surveys, household income surveys, and other records, harmonised by ILO for consistency. The dataset is in tabular format and suitable for tasks such as tabular classification, regression, and time-series forecasting.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-deap-sex-edu-geo-rt-unemployment-rate-by-sex-education-and-rural-urban 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT核心统计数据库,通过调用其REST API接口直接获取原始指标数据。数据经过筛选,仅保留欧洲39个国家的观测记录,并依据ILO劳动统计学家国际会议(ICLS)定义对原始调查微观数据进行统一协调处理。Electric Sheep Europe团队对数据进行了重新封装,将异构的原始数据转换为标准化的表格格式,并保留了数据来源标识(source.label列)以确保可追溯性。最终形成包含76,163条观测记录、覆盖1987至2025年时间跨度的结构化数据集。
特点
本数据集聚焦于按性别、教育程度及城乡区域划分的失业率指标,提供了极其精细的社会经济维度拆解。其独特之处在于融合了空间(39个欧洲国家)、时间(近四十年)与人口特征(性别、教育、城乡)三重分析视角,每条记录均附带完整的分类标签与数据质量标记(如obs_status)。数据整合了多个国家不同调查来源的协调结果,并通过标准化字段设计(如sex、classif1、classif2)使得跨国家、跨时期的比较分析成为可能。此外,数据集还提供了详细的注释字段,用以标记方法变更、非标准分类等数据质量信息。
使用方法
用户可通过HuggingFace Datasets库的load_dataset()函数一键加载该数据集,并将其转换为Pandas DataFrame进行灵活操作。支持按国家代码(ref_area)过滤特定国家的时间序列数据,亦可针对单一指标按时间排序进行趋势分析。利用pivot_table方法,用户能够便捷地将数据重塑为国家×年份的矩阵形式,便于进行横截面或面板数据分析。数据集包含完整的脱敏分类变量(如性别、教育水平、地域类型),适用于构建分类、回归或时间序列预测模型,为劳动经济学研究提供即开即用的结构化数据基础。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)于2025年发布,经Electric Sheep Europe在HuggingFace上重新整理,聚焦欧洲39个国家1987至2025年间按性别、教育程度及城乡区域划分的失业率指标。作为ILOSTAT数据库的核心组成部分,该数据集整合了来自劳动力调查、家庭收支调查及行政记录等多种来源的微观数据,并依据国际劳工统计学家会议(ICLS)标准进行统一化处理。其核心研究问题在于揭示欧洲不同社会人口群体在劳动力市场中的结构性差异,为劳动经济学、区域发展研究及社会政策评估提供关键实证基础。凭借覆盖近四十年的长时间序列和精细的分类维度,该数据集在分析欧洲失业动态、教育回报率及城乡就业鸿沟等议题上展现出重要影响力,成为相关领域学者和政策制定者不可或缺的量化工具。
当前挑战
该数据集所解决的领域挑战主要在于跨国家、跨时期失业率指标在性别、教育及城乡维度上的精细分层与可比性缺失,以及原始调查数据因方法修订、来源变更或样本差异导致的时序断裂与统计口径不统一问题。在构建过程中,面临的挑战包括:从ILO REST API抽取海量原始数据时需处理复杂的分类编码体系与多源标注信息,确保数据清洗后保留完整的溯源列以支持方法论追溯;针对不同来源间最佳来源的选取策略需在数据完整性与一致性间取得平衡;以及将不同国家、年份间歇性发布的年度数据整合为统一时间序列时,需应对观测值缺失、状态标记不稳定等质量问题,最终形成结构规整、可直接用于机器学习的表格化数据集。
常用场景
经典使用场景
该数据集汇聚了欧洲39个国家1987年至2025年间按性别、教育水平和城乡区域划分的失业率统计数据,共计76,163条观测记录。在劳动经济学与社会科学研究领域,它常被用于构建面板数据模型,深入剖析失业率的时空演变规律,探究教育水平对就业机会的差异化影响,以及城乡劳动市场结构的异质性特征。时序预测与回归分析亦是其重要应用场景,研究者可借助该数据集训练机器学习模型,实现对未来失业趋势的精准预估。
衍生相关工作
该数据集衍生了一系列具有深远影响的经典研究工作。以面板数据为基础,学者们开发了多种时空经济计量模型,如动态空间杜宾模型,用以刻画失业率的空间溢出效应。在机器学习领域,研究者利用该数据集训练了基于梯度提升树与长短期记忆网络的失业率预测模型,显著提升了短期预测精度。另一类衍生工作聚焦于维度归约技术,通过主成分分析与聚类方法识别欧洲劳动市场的区域性模式,为后续的跨国比较研究奠定了方法基石。
数据集最近研究
最新研究方向
当前,基于ILOSTAT数据的欧洲失业率研究正深入探讨人口异质性对劳动力市场的影响,尤其聚焦性别、教育水平与城乡区域之间失业率的动态差异。这一方向与近年来欧洲区域经济不平等、教育回报率下滑及城乡发展失衡等社会热点紧密相连。该数据集提供了1987至2025年间39个欧洲国家的细分观测值,为构建多维度时间序列模型、评估政策干预效果以及揭示结构性失业的深层动因提供了坚实的数据基础。其严谨的指标定义与丰富的分类标签,有力支撑了从宏观劳动力市场到微观个体特征的精确定量分析,对推动包容性就业政策的制定与评估具有重要学术与实践意义。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务