遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-edu-cbr-nb-unemployment-by-sex-education-and-place-of-birth-t

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - international-migrant-stock - ilo - labour - employment pretty_name: "Unemployment by sex, education and place of birth (thousands) | Europe (ILOSTAT)" --- # Unemployment by sex, education and place of birth (thousands) | Europe (ILOSTAT) 🇪🇺 **33,775 observations** · **36 Europe countries** · **1987–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-33,775-blue) ![countries](https://img.shields.io/badge/countries-36-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **33,775 observations** of `International migrant stock` data across **36 Europe countries**, spanning **1987–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_CBR_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** International migrant stock ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_EDU_CBR_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 36 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GRC` | 1,840 | 1987 | 2025 | | `SWE` | 1,494 | 1995 | 2024 | | `FRA` | 1,485 | 1995 | 2024 | | `NLD` | 1,444 | 1996 | 2024 | | `GBR` | 1,442 | 1995 | 2025 | | `NOR` | 1,355 | 1996 | 2024 | | `PRT` | 1,344 | 1995 | 2025 | | `ESP` | 1,334 | 1995 | 2025 | | `BEL` | 1,307 | 1995 | 2024 | | `IRL` | 1,276 | 1999 | 2024 | | `DNK` | 1,262 | 1995 | 2024 | | `AUT` | 1,151 | 1995 | 2025 | | `LUX` | 1,114 | 1995 | 2024 | | `FIN` | 1,067 | 1995 | 2024 | | `CHE` | 1,003 | 2001 | 2025 | | ... | _21 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_EDU_CBR_NB` — Unemployment by sex, education and place of birth (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:480` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_EDU_CBR_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, education and pl…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `classif2` | `string` | Second classification variable where applicable | `CBR_BIR_TOTAL` | | `classif2.label` | `string` | — | `Place of birth: Total` | | `time` | `int64` | Observation year | `2024` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `108.247` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:5578` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-edu-cbr-nb-unemployment-by-sex-education-and-place-of-birth-t") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_EDU_CBR_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_EDU_CBR_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_EDU_CBR_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_edu_cbr_nb_unemployment_by_sex_education_and_place_of_birth_t_2025, title = {Unemployment by sex, education and place of birth (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_CBR_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-edu-cbr-nb-unemployment-by-sex-education-and-place-of-birth-t}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_CBR_NB_

license: CC BY 4.0 language: - en task_categories: - 表格分类 - 表格回归 - 时间序列预测 multilinguality: 单语言 size_categories: - 10K<n<100K tags: - 表格数据 - 欧洲 - ILOSTAT - 国际移民存量 - 国际劳工组织(ILO) - 劳工 - 就业 pretty_name: "按性别、教育程度与出生地区划分的失业人数(千人)| 欧洲(ILOSTAT)" # 按性别、教育程度与出生地区划分的失业人数(千人)| 欧洲(ILOSTAT) 🇪🇺 **33,775 条观测数据** · **36个欧洲国家** · **1987–2025年** · *由[Electric Sheep Europe](https://huggingface.co/electricsheepeurope)重新整理* ![rows](https://img.shields.io/badge/rows-33,775-blue) ![countries](https://img.shields.io/badge/countries-36-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## 快速摘要(TL;DR) 本数据集包含覆盖**36个欧洲国家**、跨度**1987–2025年**的**33,775条**`国际移民存量`(International migrant stock)观测数据,涵盖**1项独立指标**。 ## 数据源说明 **国际劳工组织统计数据库(ILOSTAT)** 是国际劳工组织(ILO)的核心统计数据库,也是全球领先的劳工统计权威来源。其收录涵盖就业、失业、薪资、工作时长、童工、非正规经济、社会保障、职业伤害以及可持续发展目标(SDG)体面工作目标等多类指标,数据来源于全国劳动力调查、家庭收入调查、机构调查及行政记录,覆盖全球200余个经济体,由国际劳工组织统计司负责数据的标准化协调。 - **数据源地址**:[ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_CBR_NB) - **发布方**:国际劳工组织(ILO) - **许可协议**:[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) - **主题**:国际移民存量 ## 数据处理方法 本数据集直接从ILOSTAT的REST API接口`https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_EDU_CBR_NB`拉取原始数据,并筛选出欧洲地区的ISO3国家编码数据。ILOSTAT采用国际劳工统计学家会议(ICLS)定义对原始调查微观数据进行标准化协调,数据来源信息将在`source.label`列中标记以保证可追溯性。 ## 地理覆盖范围 36个欧洲国家,以下为按数据行数排序的前10个国家示例: | 国家 | 数据行数 | 起始年份 | 结束年份 | |---------|-----:|-----------:|----------:| | `GRC` | 1,840 | 1987 | 2025 | | `SWE` | 1,494 | 1995 | 2024 | | `FRA` | 1,485 | 1995 | 2024 | | `NLD` | 1,444 | 1996 | 2024 | | `GBR` | 1,442 | 1995 | 2025 | | `NOR` | 1,355 | 1996 | 2024 | | `PRT` | 1,344 | 1995 | 2025 | | `ESP` | 1,334 | 1995 | 2025 | | `BEL` | 1,307 | 1995 | 2024 | | `IRL` | 1,276 | 1999 | 2024 | | `DNK` | 1,262 | 1995 | 2024 | | `AUT` | 1,151 | 1995 | 2025 | | `LUX` | 1,114 | 1995 | 2024 | | `FIN` | 1,067 | 1995 | 2024 | | `CHE` | 1,003 | 2001 | 2025 | | ... | 其余21个国家 | | | ## 指标(示例) - `UNE_TUNE_SEX_EDU_CBR_NB` — 按性别、教育程度与出生地区划分的失业人数(千人) ## 数据结构(Schema) | 列名 | 数据类型 | 说明 | 示例 | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 国家编码 | `ALB` | | `ref_area.label` | `string` | 英文国家名称 | `Albania` | | `source` | `string` | ILOSTAT 数据源编码(如劳动力调查) | `BA:480` | | `source.label` | `string` | 英文数据源名称 | `LFS - 劳动力调查` | | `indicator` | `string` | ILOSTAT 指标编码 | `UNE_TUNE_SEX_EDU_CBR_NB` | | `indicator.label` | `string` | 英文指标名称 | `Unemployment by sex, education and pl…` | | `sex` | `string` | 性别细分维度(`SEX_T`=总计,`SEX_M`=男性,`SEX_F`=女性) | `SEX_T` | | `sex.label` | `string` | 维度说明 | `Total` | | `classif1` | `string` | 第一分类变量(年龄、教育程度、身份等) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | 维度说明 | `Education (Aggregate levels): Total` | | `classif2` | `string` | 可选第二分类变量 | `CBR_BIR_TOTAL` | | `classif2.label` | `string` | 维度说明 | `Place of birth: Total` | | `time` | `int64` | 观测年份 | `2024` | | `obs_value` | `float64` | 观测指标值(单位见指标定义) | `108.247` | | `obs_status` | `string` | 观测状态标记(如临时、不可靠) | `U` | | `obs_status.label` | `string` | 状态说明 | `Unreliable` | | `note_classif` | `string` | 分类备注编码 | `C3:5578` | | `note_classif.label` | `string` | 备注说明 | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | 指标备注编码 | `I11:264` | | `note_indicator.label` | `string` | 备注说明 | `Break in series: Methodology revised` | | `note_source` | `string` | 数据源备注编码 | `R1:3513` | | `note_source.label` | `string` | 备注说明 | `Repository: ILO-STATISTICS - Micro da…` | ## 细分维度 以下列提供数据细分维度: - **`sex`**(共3个唯一值):`SEX_T`、`SEX_M`、`SEX_F` ## 数据质量与注意事项 - 本数据集为年度频率数据,部分指标另有月度或季度序列,未包含在本数据集中。 - 当同一国家×年份的同一指标存在多个数据源时,将采用国际劳工组织选定的“最优数据源”。 - 细分维度列(`sex`、`classif1`、`classif2`)仅在指标支持对应细分时才会有非空值。 ## 使用示例 python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-edu-cbr-nb-unemployment-by-sex-education-and-place-of-birth-t") df = ds["train"].to_pandas() print(df.head()) ### 筛选单一国家数据 python germany = df[df["ref_area"] == "DEU"] ### 单个指标的时间序列数据 python sample = (df[df["indicator"] == "UNE_TUNE_SEX_EDU_CBR_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_EDU_CBR_NB") ### 转换为国家×年份矩阵 python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_EDU_CBR_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ## 引用格式 bibtex @misc{europe_ilo_une_tune_sex_edu_cbr_nb_unemployment_by_sex_education_and_place_of_birth_t_2025, title = {Unemployment by sex, education and place of birth (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_CBR_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-edu-cbr-nb-unemployment-by-sex-education-and-place-of-birth-t}} } ## 许可协议 本数据集采用[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)协议发布。 原始数据版权归国际劳工组织(ILO)所有。使用本数据集时,请同时引用上述原始数据源及Electric Sheep Europe的重新整理版本。 ## 关于Electric Sheep Electric Sheep Europe是Electric Sheep项目的组成部分,该项目旨在为HuggingFace平台构建统一的、适用于机器学习的欧洲地区数据层。我们从权威开源数据源拉取数据,标准化数据结构,打包为Parquet格式,并发布为格式统一的数据集卡片,以便研究人员与开发者仅需通过`load_dataset()`即可在数秒内启动数据分析工作。 浏览完整数据集集合:[huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _数据溯源:2026-05-27通过Electric Sheep管道导入。源URL:https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_EDU_CBR_NB_

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-edu-cbr-nb-unemployment-by-sex-education-and-place-of-birth-t 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT统计数据库,通过REST API直接拉取指示器‘UNE_TUNE_SEX_EDU_CBR_NB’的原始数据,并依据ILO基于国际劳工统计学家会议(ICLS)定义的标准对调查微观数据进行协调处理。随后,数据被筛选至欧洲36个国家的ISO3代码范围,并经Electric Sheep Europe整合打包,形成统一的Parquet格式数据集。每一笔观测值均附有来源标签,确保数据可追溯。
特点
本数据集涵盖1987年至2025年间欧洲36个国家的33,775条观测记录,聚焦于按性别、教育程度和出生地划分的失业人数(单位:千)。其结构化特征包括多维度分解变量,如性别(总、男、女)、教育水平分类及出生地信息,支持精细的交叉分析。数据以年度频率发布,并标注了观测状态(如不可靠值),方便用户评估数据质量。此外,数据集采用CC-BY 4.0许可,兼具开放性与可复用性。
使用方法
用户可通过HuggingFace Datasets库的load_dataset()函数直接加载数据,并转换为Pandas DataFrame以便分析。示例包括按国家筛选(如德国)、对特定指示器进行时间序列可视化,以及通过透视表构建国家×年份矩阵。数据集适合用于劳动经济学中的面板数据回归、时间序列预测或分类任务,使用时需留意数据年度频率及‘最佳来源’选择策略,以避免多源冲突。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)统计司于2025年整理发布,并经Electric Sheep Europe重新封装,聚焦于欧洲36个国家1987至2025年间按性别、教育程度和出生地划分的失业人数(以千计)。作为ILOSTAT数据库的核心子集,该数据依托国际劳工统计学家会议(ICLS)的统一定义,整合了各国劳动力调查、家庭收支调查及行政记录等多源信息,旨在系统刻画移民劳动力市场的结构性特征。其在劳动经济学、人口迁移与公共政策研究中具有显著影响力,为探讨教育回报、性别差异与移民就业融合等核心问题提供了跨时空的标准化面板数据资源。
当前挑战
该数据集所解决的领域挑战在于填补欧洲跨国劳动力统计数据的碎片化空白,尤其针对失业与移民身份、教育背景的交叉维度,传统国际数据往往缺乏细粒度分类,难以支撑多维异质性分析。构建过程中面临多重挑战:首先,各国调查方法、问卷设计及数据采集频率差异显著,ILO需借助ICLS标准进行跨源数据协调并标注来源标记以维持可追溯性;其次,缺失值与观测状态复杂(如标记为‘不可靠’或‘方法修订’),需在保留原始标记的同时平衡数据可用性与质量;此外,年度频率限制了季节性与月度波动的研究,且部分国家早期年份数据稀疏,需依赖插补或选择性分析策略进行稳健推断。
常用场景
经典使用场景
该数据集汇集了欧洲36个国家长达近四十年的失业数据,按性别、教育程度与出生地三维度进行精细拆解,为劳动经济学、人口统计学及社会分层研究提供了珍贵的结构化面板数据。研究者可基于'obs_value'字段构建时间序列模型,分析不同子群体失业率的动态演化;亦可通过'sex'、'classif1'和'classif2'等分类变量,精准刻画移民与本土劳动者在劳动力市场上的结构性差异。数据集的年度频率与丰富的元数据注释,使其成为开展跨国比较研究、评估劳动力政策效果的理想素材。
解决学术问题
围绕劳动力市场中的结构性失业与人口异质性这一经典学术命题,该数据集助力学者实证检验人力资本理论与社会排斥理论。通过整合教育程度与出生地信息,研究者能够量化教育回报在移民与本国出生群体间的差异,揭示技能错配、歧视性就业障碍及代际流动困境。同时,时间跨度的纵贯性覆盖劳动力市场周期性波动,为探究经济衰退期边缘群体的就业脆弱性提供了实证支撑。这些工作深化了对欧洲劳动力市场分割机制的理解,并为包容性就业政策的制定提供了数据基础。
衍生相关工作
该数据集衍生出一系列具有影响力的学术与政策研究工作。其中,基于面板数据构建的失业持续时间模型,揭示了欧洲各国移民群体失业率持续高企的结构性成因;融合教育分层的随机效应回归,量化了高等教育对不同出生人口群体抵御失业风险的边际效应。部分研究进一步将数据与政治经济学框架结合,探讨移民失业率与民粹主义投票倾向之间的时空关联。此外,该数据集还被用作基准测试,检验针对多分类变量与稀疏时间序列的统计学习方法及可解释性框架的性能表现。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务