遇见数据集

electricsheepeurope/europe-ilo-une-deap-sex-edu-cbr-rt-unemployment-rate-by-sex-education-and-place-of-bi

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - international-migrant-stock - ilo - labour - employment pretty_name: "Unemployment rate by sex, education and place of birth (%) | Europe (ILOSTAT)" --- # Unemployment rate by sex, education and place of birth (%) | Europe (ILOSTAT) 🇪🇺 **33,775 observations** · **36 Europe countries** · **1987–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-33,775-blue) ![countries](https://img.shields.io/badge/countries-36-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **33,775 observations** of `International migrant stock` data across **36 Europe countries**, spanning **1987–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_EDU_CBR_RT) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** International migrant stock ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_DEAP_SEX_EDU_CBR_RT` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 36 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GRC` | 1,840 | 1987 | 2025 | | `SWE` | 1,494 | 1995 | 2024 | | `FRA` | 1,485 | 1995 | 2024 | | `NLD` | 1,444 | 1996 | 2024 | | `GBR` | 1,442 | 1995 | 2025 | | `NOR` | 1,355 | 1996 | 2024 | | `PRT` | 1,344 | 1995 | 2025 | | `ESP` | 1,334 | 1995 | 2025 | | `BEL` | 1,307 | 1995 | 2024 | | `IRL` | 1,276 | 1999 | 2024 | | `DNK` | 1,262 | 1995 | 2024 | | `AUT` | 1,151 | 1995 | 2025 | | `LUX` | 1,114 | 1995 | 2024 | | `FIN` | 1,067 | 1995 | 2024 | | `CHE` | 1,003 | 2001 | 2025 | | ... | _21 more countries_ | | | ## Indicators (sample) - `UNE_DEAP_SEX_EDU_CBR_RT` — Unemployment rate by sex, education and place of birth (%) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:480` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_DEAP_SEX_EDU_CBR_RT` | | `indicator.label` | `string` | Indicator name in English | `Unemployment rate by sex, education a…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `classif2` | `string` | Second classification variable where applicable | `CBR_BIR_TOTAL` | | `classif2.label` | `string` | — | `Place of birth: Total` | | `time` | `int64` | Observation year | `2024` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `8.431` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:5578` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-deap-sex-edu-cbr-rt-unemployment-rate-by-sex-education-and-place-of-bi") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_DEAP_SEX_EDU_CBR_RT"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_DEAP_SEX_EDU_CBR_RT") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_DEAP_SEX_EDU_CBR_RT"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_deap_sex_edu_cbr_rt_unemployment_rate_by_sex_education_and_place_of_bi_2025, title = {Unemployment rate by sex, education and place of birth (%) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_EDU_CBR_RT}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-deap-sex-edu-cbr-rt-unemployment-rate-by-sex-education-and-place-of-bi}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_EDU_CBR_RT_

This dataset, titled Unemployment rate by sex, education and place of birth (%) | Europe (ILOSTAT), contains unemployment rate statistics from the International Labour Organization (ILO) ILOSTAT database for 36 European countries spanning from 1987 to 2025. It comprises 33,775 observations, with the core indicator being UNE_DEAP_SEX_EDU_CBR_RT, which represents the unemployment rate disaggregated by sex, education level, and place of birth (percentage). The data is sourced via the ILOSTAT REST API, processed, and filtered to include only European countries. The dataset provides detailed disaggregation dimensions, including sex (total, male, female), education (aggregate levels), and place of birth. The schema includes columns such as country code, country name, data source, indicator code, indicator label, sex classification, education classification, place of birth classification, observation year, observed value, observation status flags, and relevant notes. It is designed for tabular classification, regression, and time-series forecasting tasks, aiming to offer a unified, machine-learning-ready data layer for European labor market analysis.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-deap-sex-edu-cbr-rt-unemployment-rate-by-sex-education-and-place-of-bi 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT统计数据库,通过REST API接口直接获取指标代码为UNE_DEAP_SEX_EDU_CBR_RT的原始数据,并依据欧洲ISO3国家代码进行地理范围筛选。数据涵盖36个欧洲国家,时间跨度从1987年至2025年,包含33,775条观测记录。ILOSTAT依据国际劳工统计学家会议(ICLS)定义对各国劳动力调查微观数据进行统一协调处理,并在source.label列中标注数据来源,以确保数据的可追溯性与标准化。最终由Electric Sheep Europe进行清洗、重封装并以Parquet格式发布。
特点
数据集以失业率为核心指标,按性别、教育程度和出生地三个维度进行精细分解,提供从总体到分组的多层次观测值。每条记录包含ISO国家代码、年份、观测值及状态标志,并附带来源、分类注释与指标说明等多列元数据,便于用户理解数据背景与质量。数据集为年度频率,在存在多个数据来源时采用ILO选定的“最佳来源”,确保数据一致性。同时,数据支持时间序列分析与面板数据建模,适合进行跨国家、跨时段的多维比较研究。
使用方法
用户可通过HuggingFace的datasets库直接加载数据集,使用load_dataset('electricsheepeurope/europe-ilo-une-deap-sex-edu-cbr-rt-unemployment-rate-by-sex-education-and-place-of-bi')命令即可获取训练集,并便捷地转换为pandas DataFrame进行后续分析。支持按国家代码ref_area字段过滤特定国家的数据子集,也可对单个指标按时间排序后进行时序可视化。此外,可通过pivot_table方法将数据重塑为国家×年份的矩阵形式,便于面板数据回归或聚类分析,满足经济学与劳动社会学领域的定量研究需求。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)统计司依托ILOSTAT数据库构建,经Electric Sheep Europe于2025年整合发布,旨在系统记录欧洲36个国家1987至2025年间按性别、教育程度及出生地划分的失业率。涵盖33,775条观测值,核心研究问题聚焦于劳动力市场中弱势群体的结构性失业差异,为劳动经济学、移民政策及社会分层研究提供了标准化时序框架。作为ILO全球劳动力统计体系的关键组成部分,该数据集通过统一ICLS定义与多源调查数据融合,推动了跨国比较分析的可信度,成为评估欧洲各国就业政策效果及迁移劳动力整合程度的重要实证基础。
当前挑战
该数据集所解决的领域核心挑战在于揭示教育水平与移民身份对失业率影响的非线性交互效应,尤其在多国比较中需克服统计口径异质性与缺失值偏差。构建过程中面临多源调查数据(如劳动力调查、行政记录)的标准化问题,当同一国家-年份存在多个数据来源时,ILO需通过算法遴选最优来源,但部分观测值仍标注为不可靠状态。此外,时序数据受方法论修订、教育分类差异及出生地定义变化影响,导致序列断裂(如note_indicator标注的断裂点),需谨慎处理交叉分类(性别×教育×出生地)造成的稀疏矩阵问题,这对无偏估计与模型泛化构成显著障碍。
常用场景
经典使用场景
该数据集源自国际劳工组织ILOSTAT数据库,聚焦于欧洲36个国家在1987至2025年间按性别、教育程度和出生地分类的失业率(%)指标,包含超过3.3万条观测记录。其经典使用场景在于支撑跨国、跨时域的劳动力市场结构分析与比较研究。研究者可借此精确解构不同社会人口群体(如本土出生与外国出生人群、不同教育层次及性别)间的失业率差异,并追踪其在长期经济周期与政策变迁下的演变轨迹。该数据集特别适合于时间序列建模和面板数据分析,是评估欧洲各国人力资本配置效率及社会融合状况的宝贵资源。
衍生相关工作
基于该数据集,已衍生出一系列影响深远的经典工作。在方法论层面,许多研究利用其多维分层特性,发展了针对不平衡面板数据的机器学习预测模型,用以捕获失业率的非线性动因。在社会政策评价方面,学者们基于该数据构造了欧洲社会包容性指数,量化了不同国家在移民劳动市场融合上的进展。特别值得关注的是,它启发了关于教育层级作为社会流动缓冲垫的跨国比较研究,揭示了高等教育在缓解移民群体失业冲击中的保护性作用,这些工作进一步巩固了ILOSTAT作为劳动力实证研究黄金标准数据源的地位。
数据集最近研究
最新研究方向
基于ILOSTAT标准化框架,该数据集为欧洲劳动力市场中的结构性失业与移民融合议题提供了精细化的量化分析基础。当前前沿研究聚焦于利用性别、教育水平与出生地三重维度,结合1987至2025年间涵盖36国的时间序列,揭示不同群体在就业机会获取上的不平等模式及周期性波动。尤其在欧洲移民危机与后疫情时代就业复苏背景下,这一多变量分层数据成为评估社会融入政策成效、识别教育回报差异以及构建劳动参与预测模型的关键素材。其与ILO国际劳工统计规范的高度契合,使得跨国的比较研究与劳动力市场异质性分析得以在方法论上实现严谨推导,对推动包容性增长与循证决策具有深远意义。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务