遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-age-geo-nb-unemployment-by-sex-age-and-rural-urban-areas-thou

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 100K<n<1M tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment by sex, age and rural / urban areas (thousands) | Europe (ILOSTAT)" --- # Unemployment by sex, age and rural / urban areas (thousands) | Europe (ILOSTAT) 🇪🇺 **122,942 observations** · **40 Europe countries** · **1987–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-122,942-blue) ![countries](https://img.shields.io/badge/countries-40-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **122,942 observations** of `Unemployment` data across **40 Europe countries**, spanning **1987–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_GEO_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_AGE_GEO_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 40 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GRC` | 5,430 | 1987 | 2025 | | `ITA` | 4,802 | 1992 | 2025 | | `DEU` | 4,615 | 1992 | 2025 | | `FRA` | 4,606 | 1992 | 2025 | | `PRT` | 4,591 | 1992 | 2025 | | `ESP` | 4,570 | 1992 | 2025 | | `DNK` | 4,480 | 1992 | 2025 | | `IRL` | 4,435 | 1993 | 2025 | | `NLD` | 4,398 | 1992 | 2025 | | `BEL` | 4,254 | 1992 | 2025 | | `SWE` | 4,252 | 1995 | 2025 | | `GBR` | 4,173 | 1992 | 2019 | | `AUT` | 4,084 | 1995 | 2025 | | `FIN` | 3,976 | 1995 | 2025 | | `LTU` | 3,674 | 1998 | 2025 | | ... | _25 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_AGE_GEO_NB` — Unemployment by sex, age and rural / urban areas (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:7401` | | `source.label` | `string` | Source name in English | `HIES - Living Standards Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_AGE_GEO_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, age and rural / …` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHADULT_YGE15` | | `classif1.label` | `string` | — | `Age (Youth, adults): 15+` | | `classif2` | `string` | Second classification variable where applicable | `GEO_COV_NAT` | | `classif2.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2012` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `207.786` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C12:2951` | | `note_classif.label` | `string` | — | `Rural / urban areas coverage definiti…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-age-geo-nb-unemployment-by-sex-age-and-rural-urban-areas-thou") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_GEO_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_AGE_GEO_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_GEO_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_age_geo_nb_unemployment_by_sex_age_and_rural_urban_areas_thou_2025, title = {Unemployment by sex, age and rural / urban areas (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_GEO_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-age-geo-nb-unemployment-by-sex-age-and-rural-urban-areas-thou}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_GEO_NB_

This dataset contains 122,942 observations of unemployment data across 40 Europe countries, spanning from 1987 to 2025, focusing on one distinct indicator: Unemployment by sex, age and rural/urban areas (in thousands). The data is sourced from the International Labour Organizations ILOSTAT database, retrieved via API and filtered to European ISO3 country codes. It includes disaggregation dimensions such as sex (total, male, female), age groups, and rural/urban area types, with columns for country codes, indicator labels, source information, observation values, and quality flags. The dataset is structured in tabular format, suitable for tasks like tabular classification, regression, and time-series forecasting. Data is annual and harmonized using ILOs International Conference of Labour Statisticians definitions for consistency.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-age-geo-nb-unemployment-by-sex-age-and-rural-urban-areas-thou 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的核心统计数据库ILOSTAT,聚焦于欧洲地区失业状况的精细刻画。通过调用ILOSTAT REST API,获取了标识为UNE_TUNE_SEX_AGE_GEO_NB的原始失业指标数据,并依据欧洲ISO3国家代码进行地理范围筛选。所有数据均经过ILO基于国际劳工统计学家会议(ICLS)定义的标准方法进行统一整合与质量控制,原始调查微观数据的来源信息亦被完整保留于' source.label'字段中,确保了数据溯源的可信度。
使用方法
用户可通过HuggingFace Datasets库便捷加载该数据集,仅需执行`load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-age-geo-nb-unemployment-by-sex-age-and-rural-urban-areas-thou")`命令即可获取训练集,并进一步转换为Pandas DataFrame进行分析。典型应用场景包括按国家过滤数据以开展国别研究,或针对单一指标进行时间序列可视化以揭示失业率演变趋势。此外,用户还可利用数据透视功能构建以年份为行、国家为列的矩阵表格,为跨区域比较与计量建模奠定基础。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)于2025年通过其ILOSTAT数据库发布,并由Electric Sheep Europe团队重新整理封装。核心研究问题在于系统记录欧洲40个国家自1987年至2025年间,按性别、年龄及城乡地域划分的失业人口数据(以千人为单位)。作为全球劳动统计的权威来源,ILOSTAT基于国际劳工统计学家会议标准,融合了劳动力调查、住户收支调查及行政记录等多源数据,为劳动经济学、社会保障政策及可持续发展目标监测提供了关键支撑。该数据集凭借其长期时间跨度与精细分层维度,对欧洲劳动力市场的时空演变研究、城乡失业差异分析及性别平等评估具有显著影响力。
当前挑战
该数据集所应对的领域挑战主要源于失业统计的复杂异构性:各国调查口径、城乡界定标准及年龄分段方式存在差异,导致跨国家与跨时期的数据可比性受限。构建过程中面临的挑战包括多源数据融合的归一化难题,例如需处理因方法修订导致的时间序列断裂(如注释中的‘Break in series’),以及‘最佳来源’选取规则对数据一致性的影响。此外,数据质量标注的多样性(如‘不可靠’、‘临时性’)增加了建模时的不确定性,而城乡分类(如‘GEO_COV_NAT’)的覆盖范围定义差异进一步加剧了空间维度的分析难度。这些挑战要求研究者必须审慎处理数据溯源标记与分类变量中的空值风险。
常用场景
经典使用场景
在欧洲劳动力市场研究中,该数据集因其细粒度的失业统计指标而成为不可或缺的基石。它提供了覆盖40个欧洲国家、跨越近四十年(1987-2025年)的失业数据,并按照性别、年龄以及城乡区域进行了细致划分,共计超过12万条观测记录。研究者可以借此构建面板数据模型,剖析不同人口亚群在时间维度上的失业动态,或运用时间序列分析方法,探究欧洲各国失业率的长期趋势与周期性波动,从而为理解劳动力市场的结构变迁提供实证支撑。
解决学术问题
该数据集致力于解决欧洲范围内失业问题研究中数据碎片化与口径不统一的核心症结。通过整合ILOSTAT这一权威来源并统一分类标准,它使学者能够跨越国界进行严谨的比较分析。具体而言,它助力探究性别与年龄结构对失业风险的差异化影响,量化城乡二元经济结构下失业率的空间异质性,以及评估宏观经济政策或技术变革对不同群体就业的冲击。其深远意义在于为劳动经济学、人口经济学及区域科学等领域提供了高信效度的实证基础,推动了针对欧洲劳动力市场失灵与包容性增长策略的理论深化。
实际应用
在实际应用领域,该数据集的价值尤为显著,它直接服务于欧洲各国政府及国际组织的就业政策制定与效果评估。政策分析师可以利用其中按年龄和性别细分的城乡失业数据,精准识别最脆弱的就业群体,从而设计更具靶向性的职业培训和就业援助项目。同时,在商业层面,跨国企业的投资决策者能够通过分析不同地区与人群的失业趋势,评估区域市场的人力资源储备与消费潜力,优化运营布局。该数据集也为人道主义组织评估社会融合风险、监测可持续发展目标(SDG)中的体面工作进展提供了量化工具。
数据集最近研究
最新研究方向
在全球劳动力市场波动加剧与欧洲绿色转型、数字变革交织的背景下,ILOSTAT整合的欧洲失业数据集为前沿研究提供了重要支撑。当前热点方向包括利用该数据集构建高精度时序预测模型,结合性别、年龄及城乡地理维度,解析结构性失业与周期性失业的深层动因;同时,借助其长跨期面板数据(1987-2025)与标准化分类体系,学者们正深入探索欧盟就业政策冲击、新冠疫情期间劳动力市场恢复的非对称性,以及人工智能对青年、女性和农村群体就业替代效应的量化评估。该数据集的完整性与跨域可比性,为评估欧洲“技能公约”与“公正转型”机制的实际成效提供了实证基石,推动了机器学习在劳动力经济学中从描述性统计向因果推断的范式跃迁。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务