遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-age-edu-nb-unemployment-by-sex-age-and-education-thousands

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 100K<n<1M tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment by sex, age and education (thousands) | Europe (ILOSTAT)" --- # Unemployment by sex, age and education (thousands) | Europe (ILOSTAT) 🇪🇺 **419,089 observations** · **41 Europe countries** · **1976–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-419,089-blue) ![countries](https://img.shields.io/badge/countries-41-green) ![years](https://img.shields.io/badge/years-1976–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **419,089 observations** of `Unemployment` data across **41 Europe countries**, spanning **1976–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_EDU_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_AGE_EDU_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 41 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GRC` | 22,957 | 1987 | 2025 | | `GBR` | 17,215 | 1987 | 2025 | | `CHE` | 15,449 | 1991 | 2025 | | `FRA` | 15,367 | 1993 | 2024 | | `ITA` | 14,056 | 1992 | 2024 | | `CZE` | 13,880 | 1991 | 2024 | | `IRL` | 13,541 | 1988 | 2024 | | `SWE` | 13,411 | 1987 | 2024 | | `PRT` | 13,381 | 1992 | 2025 | | `ESP` | 13,019 | 1976 | 2025 | | `DEU` | 12,693 | 1992 | 2024 | | `NLD` | 12,419 | 1996 | 2024 | | `BEL` | 12,168 | 1992 | 2024 | | `AUT` | 11,875 | 1985 | 2025 | | `DNK` | 11,871 | 1992 | 2024 | | ... | _26 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_AGE_EDU_NB` — Unemployment by sex, age and education (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:480` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_AGE_EDU_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, age and educatio…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHADULT_YGE15` | | `classif1.label` | `string` | — | `Age (Youth, adults): 15+` | | `classif2` | `string` | Second classification variable where applicable | `EDU_AGGREGATE_TOTAL` | | `classif2.label` | `string` | — | `Education (Aggregate levels): Total` | | `time` | `int64` | Observation year | `2024` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `108.247` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:2138` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-age-edu-nb-unemployment-by-sex-age-and-education-thousands") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_EDU_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_AGE_EDU_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_EDU_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_age_edu_nb_unemployment_by_sex_age_and_education_thousands_2025, title = {Unemployment by sex, age and education (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_EDU_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-age-edu-nb-unemployment-by-sex-age-and-education-thousands}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_EDU_NB_

This dataset contains unemployment statistics sourced from the International Labour Organization (ILO) ILOSTAT database, focusing on the European region. It includes 419,089 observations covering 41 European countries over the period from 1976 to 2025. The core indicator of the dataset is UNE_TUNE_SEX_AGE_EDU_NB, which denotes the number of unemployed people (in thousands) categorized by sex, age and educational attainment. The data was acquired via the ILOSTAT REST API and processed to exclusively include European countries. Organized in a tabular format, the dataset features columns such as country code, gender category, age and educational category, year, and observed value, among others. It supports tasks including tabular classification, regression and time series forecasting. The data is of annual frequency, and includes data source and quality annotations, making it applicable for labor market analysis and machine learning applications.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-age-edu-nb-unemployment-by-sex-age-and-education-thousands 数据集图片
构建方式
该数据集源于国际劳工组织(ILO)的ILOSTAT中央统计数据库,聚焦于欧洲41个国家1976至2025年间按性别、年龄及教育程度划分的失业人数指标。数据通过直接调用ILOSTAT REST API接口获取原始观测值,并依据欧洲ISO3国家代码进行地理范围过滤。为确保统计口径的全球可比性,ILO依据国际劳工统计学家会议(ICLS)定义对原始调查微观数据进行了系统化协调与整合,最终形成包含419,089条观测记录的标准化表格数据,涵盖性别、年龄组别及教育层次等多维分类变量。
使用方法
该数据集以HuggingFace Datasets标准格式发布,用户可通过`load_dataset()`函数一行加载为Pandas DataFrame,即刻展开交互式分析。典型使用场景包括:按国家筛选(如`df[df['ref_area'] == 'DEU']`)获取德国子集;按时间排序绘制特定指标的时间序列折线图;以及通过数据透视表构建国家×年份的观测值矩阵,以比较不同国家间失业模式的异同。数据集设计面向机器学习的表格回归、分类以及时间序列预测任务,兼容常用的Python数据科学生态系统。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)统计司基于其核心数据库ILOSTAT构建,并由Electric Sheep Europe于2025年重新整理发布,聚焦欧洲41个国家1976至2025年间按性别、年龄与教育程度划分的失业人数(单位:千)。作为全球劳动统计的权威来源,ILOSTAT整合了各国劳动力调查、住户收入调查及行政记录等多元数据,经国际劳工统计学家会议(ICLS)定义统一标准化,为研究欧洲劳动力市场结构性变迁提供了高分辨率的时间序列观测值,共计逾41.9万条记录。该数据集在劳动经济学、社会政策分析及机器学习时序预测领域具有显著应用价值,尤其支撑了关于教育回报、性别就业差异及代际失业风险的实证研究,推动了跨国比较分析的标准化进程。
当前挑战
该领域面临的核心挑战在于跨国家、跨时期数据的一致性维护,例如各国调查方法修订、教育分类体系差异(如非标准教育层级标识)及间断序列的处理,数据集通过断点标记(Break in series)与数据源追溯字段加以应对。构建过程中,需从ILOSTAT REST API抽取原始数据,过滤至欧洲ISO3国家代码,并对同一国家-年份组合中的多来源数据依ILO首选源规则进行择优,同时确保按性别、年龄、教育程度三个维度细分的缺失值标注透明。此外,年度频率与部分指标存在的月/季频缺失,形成了时间序列建模中的粒度权衡,要求研究者审慎处理观测状态标志(如不可靠值)以保障结论稳健性。
常用场景
经典使用场景
该数据集汇聚了国际劳工组织(ILO)权威发布的ILOSTAT统计数据,涵盖了1976年至2025年间41个欧洲国家按性别、年龄和教育程度细分的失业人数(单位:千)。研究者常将其作为欧洲劳动力市场结构分析的基石,通过时间序列建模与面板数据回归,揭示失业率在不同人口群体和学历层次间的动态演化规律。其丰富的分类变量(如性别、年龄组、教育水平)为交叉分析提供了优雅的入口,可精准刻画青年失业、女性就业困境以及高等教育对就业缓冲效应的时空异质性。
解决学术问题
该数据集有效回应了劳动经济学中关于结构性失业与教育回报率的经典悖论。通过整合长达半个世纪的跨国年度观测值,学者得以剥离周期性与摩擦性失业的干扰,系统检验人力资本理论在统一劳动力市场中的适用性。它尤其助力破解“高学历未必高就业”的欧洲迷思,为最低工资制度、技能培训政策与移民涌入对本土就业的冲击提供量化依据。数据的标准化分类(如ICLS定义)确保了跨国比较的可信度,推动了计量方法在非平稳面板数据中的稳健性革新。
实际应用
在公共政策与商业决策领域,该数据集是动态劳动力监测的核心工具。各国劳工部门可据此追踪特定教育层级的失业动态,快速识别技能供需错配的行业与地区,从而优化职业培训预算的分配。欧洲央行与投资机构利用其细粒度分类,构建区域就业风险指数,辅助货币政策制定与资产配置。人力资源科技企业则将其接入薪资预测与人才流动模型,通过历史与实时的失业剖面相映射,提升校园招聘与再就业服务的精准触达效率。
数据集最近研究
最新研究方向
该数据集聚焦于欧洲劳动力市场中失业状况的精细化解构,其前沿研究方向体现在基于性别、年龄和教育程度三重维度的交叉分析上。通过整合国际劳工组织ILOSTAT自1976年至2025年间覆盖41个欧洲国家的近42万条观测记录,研究热点集中于剖析后疫情时代欧洲劳动力市场的结构性转型、高学历群体与青年失业率的动态演变,以及性别就业差异在人工智能与绿色经济转型背景下的重塑。这一数据资源为评估欧洲各国劳动政策效果、构建预测性时间序列模型提供了坚实支撑,其深远意义在于推动劳动力市场研究从宏观总量分析向微观群体画像过渡,助力精准施策与可持续就业目标的实现。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务