遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-dur-geo-nb-unemployment-by-sex-duration-and-rural-urban-areas

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment by sex, duration and rural / urban areas (thousands) | Europe (ILOSTAT)" --- # Unemployment by sex, duration and rural / urban areas (thousands) | Europe (ILOSTAT) 🇪🇺 **33,970 observations** · **37 Europe countries** · **1987–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-33,970-blue) ![countries](https://img.shields.io/badge/countries-37-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **33,970 observations** of `Unemployment` data across **37 Europe countries**, spanning **1987–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_DUR_GEO_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_DUR_GEO_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 37 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `ITA` | 1,482 | 1992 | 2024 | | `FRA` | 1,472 | 1992 | 2024 | | `IRL` | 1,432 | 1993 | 2024 | | `DEU` | 1,425 | 1992 | 2024 | | `GRC` | 1,400 | 1987 | 2024 | | `NLD` | 1,386 | 1992 | 2024 | | `DNK` | 1,308 | 1992 | 2024 | | `PRT` | 1,293 | 1992 | 2024 | | `BEL` | 1,269 | 1992 | 2024 | | `ESP` | 1,239 | 1992 | 2024 | | `FIN` | 1,228 | 1995 | 2024 | | `GBR` | 1,220 | 1992 | 2019 | | `AUT` | 1,190 | 1995 | 2025 | | `SWE` | 1,190 | 1995 | 2024 | | `MDA` | 1,080 | 2000 | 2025 | | ... | _22 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_DUR_GEO_NB` — Unemployment by sex, duration and rural / urban areas (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:7401` | | `source.label` | `string` | Source name in English | `HIES - Living Standards Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_DUR_GEO_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, duration and rur…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `DUR_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Duration (Aggregate): Total` | | `classif2` | `string` | Second classification variable where applicable | `GEO_COV_NAT` | | `classif2.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2012` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `207.786` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C7:2844` | | `note_classif.label` | `string` | — | `Nonstandard duration of unemployment:…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-dur-geo-nb-unemployment-by-sex-duration-and-rural-urban-areas") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_DUR_GEO_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_DUR_GEO_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_DUR_GEO_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_dur_geo_nb_unemployment_by_sex_duration_and_rural_urban_areas_2025, title = {Unemployment by sex, duration and rural / urban areas (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_DUR_GEO_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-dur-geo-nb-unemployment-by-sex-duration-and-rural-urban-areas}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_DUR_GEO_NB_

This dataset contains 33,970 observations of unemployment data across 37 European countries, spanning from 1987 to 2025, covering one distinct indicator: Unemployment by sex, duration and rural/urban areas (in thousands). The data is sourced from the ILOSTAT database of the International Labour Organization (ILO), retrieved via API and harmonized using ICLS definitions for consistency. It includes dimensions such as country codes, years, sex (total, male, female), classification variables (e.g., unemployment duration, area type), observed values, and status flags. The data is provided at an annual frequency and is suitable for tasks like tabular classification, regression, and time-series forecasting. Repackaged by Electric Sheep Europe under the cc-by-4.0 license, it is designed for easy integration into machine learning workflows.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-dur-geo-nb-unemployment-by-sex-duration-and-rural-urban-areas 数据集图片
构建方式
该数据集源于国际劳工组织(ILO)的核心统计数据库ILOSTAT,通过其REST API直接抽取原始指标数据,并聚焦于欧洲37个国家的失业状况。构建过程中,研究人员将数据过滤至欧洲ISO3国家代码,并依据国际劳工统计学家会议(ICLS)的定义对原始调查微观数据进行统一协调与标准化处理。每个观测值均附有来源标识,确保了数据的可追溯性与透明度。最终形成涵盖1987至2025年、包含33970条观测记录的表格型数据集,并以Parquet格式封装,便于机器学习应用。
特点
本数据集的核心特色在于其精细的多维分类结构,不仅按性别(总、男性、女性)对失业数据进行拆解,还融合了失业持续时间与城乡地域的划分,从而在单一指标下提供了丰富的分析层次。其覆盖范围横跨37个欧洲国家,时间跨度长达近四十年,为纵向比较与面板数据分析奠定了坚实基础。此外,每条记录均包含观测值状态标记及详尽的注释信息,如实反映数据修订、方法变更及可靠性问题,极大增强了数据的科学严谨性与实用性。
使用方法
使用者可通过HuggingFace的`datasets`库轻松加载该数据集,一句`load_dataset`命令即可将数据转换为Pandas DataFrame格式进行探索。针对时间序列分析,可依据国家代码对特定国家的失业趋势进行筛选与可视化。研究人员还能利用指示器代码进行数据透视,构建国家-年份矩阵,便于进行跨国的面板数据回归或聚类分析。该数据集与常见的表格分类、回归及时间序列预测任务高度兼容,为欧洲劳动力市场研究提供了即取即用的高质量数据源。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)统计部门创建,并经由Electric Sheep Europe于2025年重新整合发布,旨在提供欧洲37个国家1987至2025年间按性别、失业时长及城乡区域划分的失业人数(单位:千)的年度观测数据,共包含33,970条记录。作为ILOSTAT这一全球劳动统计核心数据库的重要组成部分,该数据集为研究欧洲劳动力市场结构变迁、性别就业差异及区域经济分化提供了标准化的时间序列基础,在劳动经济学、公共政策评估及可持续发展目标(SDG)监测领域具有广泛的应用价值。
当前挑战
该数据集所应对的核心领域挑战在于,欧洲各国失业统计口径(如ICLS定义)、调查方法(劳动力调查与行政记录并存)及数据质量参差不齐,导致跨国比较与趋势分析常因异质性而产生偏差。构建过程中面临的主要挑战包括:多源数据(如各国劳动力调查、家庭收支调查)的格式与标识符不统一,需通过ILO的‘最佳来源’筛选规则进行整合;缺失值、断点(如方法论修订导致的序列不连续)及可靠性标志(如'U'类标记)的处理,要求研究者审慎评估数据质量;此外,长期时间序列中城乡分类与失业时长定义的变化,进一步增加了跨期分析的复杂性。
常用场景
经典使用场景
在欧洲劳动力市场研究中,该数据集被广泛用于分析失业现象的性别差异、持续时间分布以及城乡异质性。基于ILOSTAT提供的统一统计口径,研究者能够通过时间序列建模与面板数据分析,探究欧洲37国自1987年至2025年间失业率的演变规律。经典使用方式包括构建失业人口按性别和城乡分组的可视化趋势图、建立固定效应或随机效应模型以识别不同国家失业持续时间的决定因素,以及利用该数据集训练分类或回归模型,预测特定社会经济条件下的失业水平。此外,该数据集的年度频率和豐富的元数据使其成为宏观经济指标协同分析的重要基础。
解决学术问题
该数据集有效解决了欧洲劳动力市场中数据碎片化与统计口径不一致的学术难题。通过融合多国长期观测值并标准化至ILO国际劳工统计标准,它使得跨国家、跨时期的失业结构比较成为可能。学术界借助该数据可深入探讨性别差异对失业持续时间的影响、城乡二元经济结构下的就业脆弱性,以及劳动市场政策在降低结构性失业中的效果。该数据集为验证劳动经济学中的搜寻与匹配理论、人力资本贬值假说提供了实证土壤,其意义在于推动欧洲层面劳动政策的精准制定与评估,尤其为可持续发展目标中有关体面劳动与经济增长的量化研究贡献了关键数据支持。
衍生相关工作
该数据集的衍生工作涵盖了多个前沿研究方向。在计量经济学领域,学者基于该数据发展了面板协整分析方法,揭示失业持续时间与经济增长的非线性关系。在机器学习范畴,研究者构造了以国家、性别、时间及城乡类别为特征的多任务预测模型,用以推断短期失业率变动。部分工作聚焦于数据可视化,开发了交互式仪表盘以实时监控欧洲失业格局的时空演化。还有研究整合该数据与劳工调查微观数据,生成合成面板,用于分析个体层面失业经历的动态特征。这些衍生产品不仅丰富了劳动统计的方法论工具,也极大提升了公开数据在政策科学中的利用效率。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务