遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-edu-geo-nb-unemployment-by-sex-education-and-rural-urban-area

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment by sex, education and rural / urban areas (thousands) | Europe (ILOSTAT)" --- # Unemployment by sex, education and rural / urban areas (thousands) | Europe (ILOSTAT) 🇪🇺 **76,163 observations** · **39 Europe countries** · **1987–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-76,163-blue) ![countries](https://img.shields.io/badge/countries-39-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **76,163 observations** of `Unemployment` data across **39 Europe countries**, spanning **1987–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_GEO_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_EDU_GEO_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 39 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GRC` | 4,805 | 1987 | 2025 | | `FRA` | 3,128 | 1993 | 2024 | | `ITA` | 2,895 | 1992 | 2024 | | `IRL` | 2,886 | 1993 | 2024 | | `SWE` | 2,776 | 1995 | 2024 | | `PRT` | 2,774 | 1992 | 2024 | | `ESP` | 2,748 | 1992 | 2024 | | `DEU` | 2,677 | 1992 | 2024 | | `BEL` | 2,647 | 1992 | 2024 | | `DNK` | 2,611 | 1992 | 2024 | | `NLD` | 2,592 | 1996 | 2024 | | `AUT` | 2,574 | 1995 | 2025 | | `SRB` | 2,370 | 2007 | 2025 | | `HUN` | 2,236 | 2001 | 2024 | | `FIN` | 2,131 | 1995 | 2024 | | ... | _24 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_EDU_GEO_NB` — Unemployment by sex, education and rural / urban areas (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:7401` | | `source.label` | `string` | Source name in English | `HIES - Living Standards Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_EDU_GEO_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, education and ru…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `classif2` | `string` | Second classification variable where applicable | `GEO_COV_NAT` | | `classif2.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2012` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `207.786` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:5578` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-edu-geo-nb-unemployment-by-sex-education-and-rural-urban-area") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_EDU_GEO_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_EDU_GEO_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_EDU_GEO_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_edu_geo_nb_unemployment_by_sex_education_and_rural_urban_area_2025, title = {Unemployment by sex, education and rural / urban areas (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_GEO_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-edu-geo-nb-unemployment-by-sex-education-and-rural-urban-area}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_GEO_NB_

This dataset contains 76,163 observations of unemployment data across 39 Europe countries, spanning 1987–2025, covering 1 distinct indicator: Unemployment by sex, education and rural / urban areas (thousands). It is sourced from the International Labour Organizations ILOSTAT database, retrieved via REST API and filtered to Europe countries, with annual frequency, and includes detailed columns such as country, source, indicator, sex classification, education classification, rural/urban classification, time, and observed values.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-edu-geo-nb-unemployment-by-sex-education-and-rural-urban-area 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT中央统计数据库中关于欧洲地区失业率的官方指标(UNE_TUNE_SEX_EDU_GEO_NB),通过直接调用ILOSTAT REST API获取原始数据,并依据国际劳工统计学家会议(ICLS)定义对各国劳动力调查、家庭收入调查等微观数据进行统一协调与清洗。随后,数据被过滤至39个欧洲国家的ISO3国家代码范围,并按照性别、教育水平及城乡地域三个维度进行分层聚合,最终打包为76,163条观测记录的表格数据集,经由Electric Sheep Europe团队重新封装发布。
特点
该数据集的核心特色在于其多维度的细粒度分层结构,涵盖性别(男性、女性、总计)、教育水平(总量与细分级别)以及城乡地域(国家、城市、乡村)三大分类变量,使用户能够深入剖析欧洲不同群体的失业构成差异。数据时间跨度覆盖1987年至2025年近40年,包含39个欧洲国家,具有年度频率的长期纵向追踪能力。所有观测值均携带来源标记、观测状态标志及详细的注释说明(如方法修订与数据断点),确保数据的透明性与可追溯性。
使用方法
用户可通过HuggingFace Datasets库的load_dataset函数直接加载该数据集,返回的“train”分割数据可便捷地转换为Pandas DataFrame进行后续分析。针对单一国家(如“DEU”)的筛选、特定指标的时间序列可视化,以及构建国家-年份透视矩阵等操作均有官方示例代码支持。数据集的模式预定义了ref_area(国家代码)、sex(性别)、classif1(教育水平)、classif2(地域类型)等关键列,便于直接用于表格分类、回归或时间序列预测等机器学习任务。
背景与挑战
背景概述
失业率作为衡量劳动力市场健康程度的核心指标,其时空异质性分析对社会经济政策制定至关重要。该数据集由国际劳工组织(ILO)统计部门基于ILOSTAT数据库构建,并经Electric Sheep Europe于2025年重新封装,覆盖1987至2025年间39个欧洲国家的76,163条观测记录,聚焦于按性别、教育程度及城乡地域维度划分的失业人数(千计)。数据集依托ILO在全球200多个经济体中的劳动力调查、家庭收支调查与行政记录等统一数据源,采用国际劳工统计学家会议(ICLS)标准进行微观数据协调,旨在为欧洲区域内的劳动经济学研究提供精细化、可复现的时空序列基础。其发布显著推动了跨国比较失业研究的数据可获取性,尤其在性别平等、教育回报与城镇化进程等交叉议题上,构成了政策评估与学术建模的关键支撑。
当前挑战
该数据集所应对的核心领域挑战在于:传统的失业统计往往仅提供国家层面的汇总数值,无法揭示不同性别、教育层次及城乡区位内的结构性差异,导致劳动力市场干预措施缺乏精准靶向。通过引入性别(sex)、教育聚合等级(classif1)与地域覆盖类型(classif2)三重解聚维度,数据集尝试弥合宏观指标与微观异质性之间的鸿沟。然而,构建过程中面临严峻挑战:其一,各国数据来源本身存在方法学差异与时间序列断点(如指标注释中标记的'Break in series: Methodology revised'),需依赖ILO的'最佳来源'选择策略以保证一致性;其二,缺失值与非标准教育分类的标注(如note_classif字段)增加了跨国家对照难度;其三,观测状态标识(如provisional或unreliable)的引入虽提升透明度,却迫使研究者必须处理低质量子样本带来的估计偏差风险。
常用场景
经典使用场景
该数据集收录了自1987年至2025年间欧洲39个国家的失业率统计数据,涵盖性别、教育程度及城乡区域三个关键维度的交叉分类。研究者可通过该数据集构建多维度的分层统计模型,探索不同社会群体在劳动力市场中的结构性差异。其经典使用场景包括:评估教育水平对男性和女性失业风险的差异化影响,比较城乡地理区位对就业机会的约束效应,以及追踪欧洲各国失业率随经济周期的演变轨迹。借助ILOSTAT标准化的数据收集与整合流程,学者们能够跨越国别统计口径的障碍,开展跨国比较研究,从而揭示欧洲劳动市场中性别平等、教育回报与区域发展不平衡等深层次议题。
解决学术问题
在学术研究领域,该数据集主要解决了因统计口径不一而导致欧洲劳动市场跨国比较研究的困境。它使得学者能够系统性地分析教育程度如何调节性别与地理区位对失业风险的影响,进而验证人力资本理论、劳动力市场分割理论及空间不平等假说在欧洲语境下的适用性。此外,该数据为评估欧盟及各国就业政策的干预效果提供了可靠的时间序列证据,支持学者开展因果推断和面板数据分析。其对ILO标准统计定义的严格遵循,确保了研究结论的可重复性和可比性,从而推动了劳动经济学、教育经济学与区域科学交叉领域的知识积累与理论深化。
衍生相关工作
围绕该数据集已衍生出一系列具有影响力的学术工作。研究者基于其丰富的维度划分,构建了多层线性模型与空间面板模型,深入探讨了性别工资差距与教育回报率在不同城乡背景下的异质性表现。部分文献利用该数据的时间序列特征,结合差分广义矩估计方法,评估了欧洲积极劳动力市场政策对失业持续期的长期影响。另有学者将其与欧盟统计局的其他社会经济指标(如GDP、人口迁徙率)进行融合,开发出综合性的区域就业脆弱性指数。这些衍生研究不仅深化了对欧洲劳动市场复杂性的理解,也为后续将机器学习方法应用于结构化劳动统计数据的创新尝试奠定了坚实基础。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务