遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-age-cbr-nb-unemployment-by-sex-age-and-place-of-birth-thousan

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - international-migrant-stock - ilo - labour - employment pretty_name: "Unemployment by sex, age and place of birth (thousands) | Europe (ILOSTAT)" --- # Unemployment by sex, age and place of birth (thousands) | Europe (ILOSTAT) 🇪🇺 **64,501 observations** · **37 Europe countries** · **1987–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-64,501-blue) ![countries](https://img.shields.io/badge/countries-37-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **64,501 observations** of `International migrant stock` data across **37 Europe countries**, spanning **1987–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_CBR_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** International migrant stock ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_AGE_CBR_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 37 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GRC` | 3,024 | 1987 | 2025 | | `SWE` | 2,781 | 1995 | 2025 | | `NLD` | 2,670 | 1995 | 2025 | | `FRA` | 2,566 | 1995 | 2025 | | `NOR` | 2,525 | 1995 | 2025 | | `GBR` | 2,475 | 1995 | 2025 | | `PRT` | 2,416 | 1995 | 2025 | | `ESP` | 2,390 | 1995 | 2025 | | `DNK` | 2,380 | 1995 | 2025 | | `AUT` | 2,344 | 1995 | 2025 | | `BEL` | 2,325 | 1995 | 2025 | | `IRL` | 2,264 | 1998 | 2025 | | `FIN` | 2,211 | 1995 | 2025 | | `LUX` | 1,998 | 1995 | 2024 | | `EST` | 1,951 | 1998 | 2025 | | ... | _22 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_AGE_CBR_NB` — Unemployment by sex, age and place of birth (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:480` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_AGE_CBR_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, age and place of…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHADULT_YGE15` | | `classif1.label` | `string` | — | `Age (Youth, adults): 15+` | | `classif2` | `string` | Second classification variable where applicable | `CBR_BIR_TOTAL` | | `classif2.label` | `string` | — | `Place of birth: Total` | | `time` | `int64` | Observation year | `2024` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `108.247` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `—` | | `note_classif.label` | `string` | — | `—` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-age-cbr-nb-unemployment-by-sex-age-and-place-of-birth-thousan") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_CBR_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_AGE_CBR_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_CBR_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_age_cbr_nb_unemployment_by_sex_age_and_place_of_birth_thousan_2025, title = {Unemployment by sex, age and place of birth (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_CBR_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-age-cbr-nb-unemployment-by-sex-age-and-place-of-birth-thousan}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_CBR_NB_

This dataset contains 64,501 observations of unemployment data by sex, age and place of birth (in thousands) across 37 European countries, spanning from 1987 to 2025, covering 1 distinct indicator. It is sourced from the International Labour Organization (ILO) ILOSTAT database, retrieved via its REST API, and repackaged in a tabular format suitable for tasks such as tabular classification, regression, and time-series forecasting. The data includes dimensions such as country, year, sex, age group, and place of birth, along with source and quality annotations.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-age-cbr-nb-unemployment-by-sex-age-and-place-of-birth-thousan 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT核心统计数据库,聚焦于欧洲37个国家1987年至2025年间按性别、年龄和出生地划分的失业人数(单位:千)。数据通过ILOSTAT REST API直接获取,原始指标代码为UNE_TUNE_SEX_AGE_CBR_NB,并依据欧洲ISO3国家代码进行地理范围过滤。ILOSTAT本身基于国际劳动统计学家会议(ICLS)定义,对各国的劳动力调查微观数据进行统一协调与标准化,数据来源在“source.label”字段中予以标注,确保了跨国数据的可比性与可追溯性。整个数据集经由Electric Sheep Europe团队重新封装,以Parquet格式发布,共计64,501条观测记录。
特点
该数据集的核心特质在于其精细的多维分类结构与严谨的数据质量标注。除基础的时间序列与地理维度外,每条记录均包含性别(总、男、女)、年龄段以及出生地(如本土出生与外籍出生)等分类变量,支持深层次的人口细分分析。数据还附带了观测状态标志(如“不可靠”)、序列断点说明以及来源注释等多重元数据字段,为用户评估数据可靠性提供了透明依据。值得注意的是,当同一国家与年份存在多个数据来源时,ILO会筛选出“最佳来源”,进一步提升了数据集的权威性。所有指标均以年频发布,覆盖了欧洲主要的劳动力市场。
使用方法
研究人员可通过HuggingFace Datasets库便捷地加载该数据集,核心命令为`load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-age-cbr-nb-unemployment-by-sex-age-and-place-of-birth-thousan")`,随后将训练集转换为Pandas DataFrame即可展开分析。典型的使用场景包括:按国家代码(ref_area)过滤以聚焦特定国家的失业趋势;针对单一指标按时间排序后绘制时序图,揭示失业率演变规律;或利用透视表功能构建国家×年份的矩阵,便于进行跨国面板数据分析或作为机器学习模型的特征输入。数据集中丰富的分类字段也支持按性别或出生地等维度进行分组聚合与比较研究。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)统计部门于2025年创建,并由Electric Sheep Europe团队在HuggingFace平台上重新封装发布。数据集聚焦于欧洲37个国家的失业状况,按性别、年龄和出生地维度细分,涵盖1987至2025年的64,501条观测记录。研究核心在于揭示劳动力市场中移民群体与本土人口的失业差异,为劳动经济学、移民政策评估及社会分层研究提供标准化数据支撑。作为ILOSTAT数据库的衍生资源,它整合了各国劳动力调查与行政记录,经ILO统一协调后具备跨国可比性,对理解欧洲地区结构性失业、代际就业不平等及外来劳动力融入问题具有重要参考价值,推动了劳动力统计领域的开放科学进程。
当前挑战
该数据集需应对以下挑战:首先,所解决的领域问题在于欧洲失业率统计长期存在国别口径差异(如调查周期、年龄分组标准),难以支撑跨国比较与时间序列分析,而该数据通过统一分类编码(如年龄分组YTHADULT_YGE15)与数据质量标记(如obs_status中的不可靠标识)部分缓解了此问题。其次,构建过程中面临多重困难,包括从ILO的REST API中提取数据时需处理多源调查的碎片化格式,对原始微观数据依据第19届国际劳动统计学家会议标准进行清洗与对齐,以及应对因方法论修订导致的序列中断(如note_indicator中的Break in series标记);此外,37个国家中部分早期数据存在缺失值或来源非最优,需依赖ILO的“最佳来源”选择策略,但该策略引入的主观性可能影响模型训练时的因果推断稳健性。
常用场景
经典使用场景
在欧洲劳动力市场研究的广阔图景中,该数据集作为国际劳工组织ILOSTAT官方统计的精选子集,提供了跨越1987至2025年间37个欧洲国家的失业人口数据,并按性别、年龄组和出生地进行了精细的维度拆分。其最经典的使用场景在于构建区域性和跨国的失业率面板数据,进而开展时间序列建模与计量经济学分析。研究者常利用其中‘obs_value’字段联合‘sex’、‘classif1’和‘classif2’等分类变量,探索移民背景、年龄结构以及性别差异对欧洲各国失业率长期演变的影响。该数据集特别适用于劳动力供给与需求匹配效率的评估、移民融合政策的因果推断,以及周期性经济波动下不同社群的脆弱性追踪,为理解欧洲大陆错综复杂的就业格局提供了坚实的数据基石。
实际应用
在实际应用维度上,该数据集为欧洲各国政府、国际组织及智库的政策制定与评估提供了实时可量化的支撑。例如,各国劳动部门可利用按出生地和性别分层的失业数据,精准识别在特定年龄段(如青年或壮年劳动力)中抵御失业风险能力较弱的外来移民群体,从而设计针对性的职业培训与就业扶持计划。欧盟委员会等跨国机构则能够借助该数据集监测劳动力市场融合进程,检验《欧洲技能议程》中关于公平就业机会的目标是否在各国落地生效。此外,数据中携带的来源标识(source.label)与观测状态标志(obs_status)允许用户追溯原始调查的可靠性,确保在诸如区域经济规划、移民配额调整或社会保障预算编制等关键决策中使用经过质量控制的权威统计信息。
衍生相关工作
该数据集衍生出的经典工作主要集中于劳动力经济学与公共政策领域。基于其提供的高维分类信息,诸多研究团队构建了预测欧洲移民失业风险的机器学习模型,例如利用梯度提升树或长短期记忆网络对以性别、年龄组、出生地及年度为特征的面板数据进行时序预测,识别结构性失业的高发人群。此外,一系列关于‘欧洲移民劳动力市场整合指数’的学术文献也依托该数据集计算了各成员国的相对表现排名。在计量方法层面,学者们衍生出了针对ILOSTAT数据断裂(break in series)标识的处理框架,如使用贝叶斯结构时间序列模型修正因方法论修订导致的序列断层,从而延长分析窗口。这些衍生工作不仅深化了对欧洲多元社会劳动力流动规律的理解,也为后续利用类似权威统计数据库开展跨学科研究树立了方法论范本。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务