遇见数据集

electricsheepeurope/europe-ilo-une-deap-sex-age-geo-rt-unemployment-rate-by-sex-age-and-rural-urban-areas

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 100K<n<1M tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment rate by sex, age and rural / urban areas (%) | Europe (ILOSTAT)" --- # Unemployment rate by sex, age and rural / urban areas (%) | Europe (ILOSTAT) 🇪🇺 **122,869 observations** · **40 Europe countries** · **1987–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-122,869-blue) ![countries](https://img.shields.io/badge/countries-40-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **122,869 observations** of `Unemployment` data across **40 Europe countries**, spanning **1987–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_AGE_GEO_RT) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_DEAP_SEX_AGE_GEO_RT` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 40 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GRC` | 5,430 | 1987 | 2025 | | `ITA` | 4,802 | 1992 | 2025 | | `DEU` | 4,615 | 1992 | 2025 | | `FRA` | 4,606 | 1992 | 2025 | | `PRT` | 4,591 | 1992 | 2025 | | `ESP` | 4,570 | 1992 | 2025 | | `IRL` | 4,493 | 1993 | 2025 | | `DNK` | 4,480 | 1992 | 2025 | | `NLD` | 4,398 | 1992 | 2025 | | `BEL` | 4,254 | 1992 | 2025 | | `SWE` | 4,252 | 1995 | 2025 | | `GBR` | 4,173 | 1992 | 2019 | | `AUT` | 4,084 | 1995 | 2025 | | `FIN` | 3,976 | 1995 | 2025 | | `LTU` | 3,674 | 1998 | 2025 | | ... | _25 more countries_ | | | ## Indicators (sample) - `UNE_DEAP_SEX_AGE_GEO_RT` — Unemployment rate by sex, age and rural / urban areas (%) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:7401` | | `source.label` | `string` | Source name in English | `HIES - Living Standards Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_DEAP_SEX_AGE_GEO_RT` | | `indicator.label` | `string` | Indicator name in English | `Unemployment rate by sex, age and rur…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHADULT_YGE15` | | `classif1.label` | `string` | — | `Age (Youth, adults): 15+` | | `classif2` | `string` | Second classification variable where applicable | `GEO_COV_NAT` | | `classif2.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2012` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `22.032` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C12:2951` | | `note_classif.label` | `string` | — | `Rural / urban areas coverage definiti…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-deap-sex-age-geo-rt-unemployment-rate-by-sex-age-and-rural-urban-areas") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_DEAP_SEX_AGE_GEO_RT"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_DEAP_SEX_AGE_GEO_RT") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_DEAP_SEX_AGE_GEO_RT"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_deap_sex_age_geo_rt_unemployment_rate_by_sex_age_and_rural_urban_areas_2025, title = {Unemployment rate by sex, age and rural / urban areas (%) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_AGE_GEO_RT}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-deap-sex-age-geo-rt-unemployment-rate-by-sex-age-and-rural-urban-areas}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_AGE_GEO_RT_

This dataset contains unemployment rate statistics for 40 European countries from 1987 to 2025, with the specific indicator Unemployment rate by sex, age and rural / urban areas (%). It comprises 122,869 observations, sourced from the International Labour Organizations (ILO) ILOSTAT database, retrieved via REST API and harmonized using International Conference of Labour Statisticians (ICLS) definitions. The data is organized in tabular format, including columns such as country code, country name, data source, indicator code, indicator name, sex disaggregation (total, male, female), age classification, area type, year, observed value, observation status, and more. The dataset supports tasks like tabular classification, tabular regression, and time-series forecasting, is monolingual (English), has a size category of 100K<n<1M, and is licensed under CC-BY-4.0. It is repackaged by Electric Sheep Europe as part of a unified, ML-ready data layer for Europe.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-deap-sex-age-geo-rt-unemployment-rate-by-sex-age-and-rural-urban-areas 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)旗下的ILOSTAT核心统计数据库,经由Electric Sheep Europe团队通过REST API接口直接抽取原始数据,并依据欧洲ISO3国家代码进行地理范围过滤与整合。在构建过程中,ILO遵循国际劳工统计学家会议(ICLS)的定义框架,对各国劳动力调查、住户收支调查及行政记录等微观数据进行标准化处理,确保不同来源的指标在概念与口径上协调一致。最终形成包含122,869条观测记录的表格型数据集,覆盖40个欧洲国家、时间跨度从1987年至2025年,并以高效的Parquet格式打包发布,便于机器学习流水线直接加载。
特点
本数据集的核心价值在于其精细的多维拆分能力。围绕“按性别、年龄和城乡划分的失业率”这一核心指标,数据集提供了性别(总、男、女)、年龄阶段(如青年与成年人)以及城乡地理覆盖类型等分类维度,允许研究者从不同角度剖析劳动力市场的结构性特征。此外,每条记录均附有来源标识、观测状态标记及方法修订注释,数据质量透明可溯。数据集的跨度长达近四十年,涵盖欧洲主要经济体与若干中小国家,为纵向比较与面板分析提供了扎实的时序基础。
使用方法
使用者可通过HuggingFace Datasets库以一行代码加载数据集,并直接转换为Pandas DataFrame进行数据探索。典型操作包括按ISO国家代码筛选特定国别数据、按时间序列排序绘制失业率变化曲线、或利用透视表功能构建国家×年份的矩阵式面板数据。数据集内嵌的详细列描述与分类标签(如sex.label、classif1.label)降低了字段理解门槛。对于回归分析、时序预测及分类任务,该数据集均能提供结构清晰、标签明确的多维特征,并遵循CC-BY 4.0许可协议,鼓励学术引用与二次分发。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)的统计数据库ILOSTAT整理,并由Electric Sheep Europe于2025年重新打包发布,聚焦欧洲40个国家1987至2025年间按性别、年龄及城乡划分的失业率指标。在劳动力经济学研究中,失业率的精细维度分解对于揭示结构性就业困境、评估劳动力市场政策的异质性影响至关重要。ILOSTAT作为全球劳动统计的权威来源,整合了各国劳动力调查、住户收入调查等多源数据,并通过国际劳工统计学家会议(ICLS)标准进行协调,确保了跨国的可比性。该数据集通过统一接口提供122,869条观测记录,覆盖了从青年到成年的年龄组别与城乡区域,为分析人口结构变迁、区域发展不均衡及性别就业差异提供了坚实的数据基础,已成为欧洲劳动力市场动态监测与比较研究的重要资源。
当前挑战
该数据集所解决的领域挑战在于,传统的失业率统计往往仅提供国家总体水平或有限的分层信息,难以反映农村与城市、不同年龄段及性别间失业风险的显著差异,从而限制了精准干预政策的制定。在构建过程中,ILOSTAT面临多源数据协调的难题,包括各国调查设计差异、指标定义不统一(如失业标准)以及时间序列上的方法修订,这些均可能导致数据口径的变动。此外,数据质量标注了“不可靠”(unreliable)等状态标识,提示部分观测值需谨慎使用;同时,由于数据仅为年度频率,限制了高频分析能力。地理覆盖上,部分国家数据年份不连续,且英国数据截止于2019年,反映出脱欧后数据获取的行政障碍。这些因素共同构成了在使用该数据集进行跨国时序比较与机器学习建模时的潜在挑战。
常用场景
经典使用场景
该数据集汇聚了欧洲40个国家自1987年至2025年间按性别、年龄及城乡区域划分的失业率统计数据,共计12万余条观测记录,为劳动经济学与区域发展研究提供了丰富的时序面板数据。研究者可基于此数据开展跨国别、跨时期的劳动力市场比较分析,例如构建固定效应模型或随机效应模型,探究性别差异、代际更替以及城市化进程对失业率的交互影响。该数据集的经典应用场景包括利用时间序列预测方法(如ARIMA或Prophet)对各国未来失业趋势进行推演,以及运用聚类分析识别失业模式相似的国家群组,为理解欧洲劳动力市场的结构性变迁奠定数据基础。
衍生相关工作
围绕该数据集,学术界衍生出了一系列富有影响力的经典工作。在计量经济学领域,有研究利用该面板数据构建了动态劳动需求模型,揭示了技术进步对欧洲不同性别劳动力替代效应的非对称性。在宏观经济学层面,学者们基于该数据集拟合了菲利普斯曲线,验证了欧元区内失业率与通胀关系的空间异质性。此外,该数据催生了多篇关于经济危机(如2008年全球金融危机、欧债危机)冲击下年轻人就业恢复力的研究论文。近年来,随着机器学习方法的普及,部分工作将随机森林或梯度提升树引入失业率预测,比较了传统时序模型与集成学习模型在区域失业率外推任务中的表现差异。
数据集最近研究
最新研究方向
该数据集聚焦于欧洲劳动力市场中失业率的多维细分,涵盖性别、年龄与城乡区域等交叉维度,为社会经济不平等与劳动政策评估提供了精细化的时间序列数据支持。在新冠疫情后全球就业格局重塑的背景下,该数据集可助力研究者追踪欧洲各国失业率的长期演变趋势,并揭示不同群体在劳动力市场中的脆弱性差异。其长达近四十年的观测跨度与多维度细分特征,为因果推断与政策模拟提供了坚实的数据基础,尤其适用于评估数字化与绿色转型对就业结构的异质性冲击。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务