遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-age-cct-nb-unemployment-by-sex-age-and-citizenship-thousands

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - international-migrant-stock - ilo - labour - employment pretty_name: "Unemployment by sex, age and citizenship (thousands) | Europe (ILOSTAT)" --- # Unemployment by sex, age and citizenship (thousands) | Europe (ILOSTAT) 🇪🇺 **66,202 observations** · **40 Europe countries** · **1991–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-66,202-blue) ![countries](https://img.shields.io/badge/countries-40-green) ![years](https://img.shields.io/badge/years-1991–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **66,202 observations** of `International migrant stock` data across **40 Europe countries**, spanning **1991–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_CCT_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** International migrant stock ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_AGE_CCT_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 40 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `SWE` | 2,718 | 1995 | 2025 | | `GRC` | 2,663 | 1992 | 2025 | | `NLD` | 2,634 | 1995 | 2025 | | `CHE` | 2,608 | 1991 | 2025 | | `DEU` | 2,596 | 1995 | 2025 | | `GBR` | 2,485 | 1995 | 2025 | | `BEL` | 2,457 | 1995 | 2025 | | `FRA` | 2,421 | 1995 | 2025 | | `NOR` | 2,384 | 1995 | 2025 | | `ESP` | 2,351 | 1995 | 2025 | | `AUT` | 2,309 | 1995 | 2025 | | `PRT` | 2,252 | 1995 | 2025 | | `DNK` | 2,214 | 1995 | 2024 | | `IRL` | 2,170 | 1998 | 2025 | | `EST` | 2,135 | 1998 | 2025 | | ... | _25 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_AGE_CCT_NB` — Unemployment by sex, age and citizenship (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:480` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_AGE_CCT_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, age and citizens…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHADULT_YGE15` | | `classif1.label` | `string` | — | `Age (Youth, adults): 15+` | | `classif2` | `string` | Second classification variable where applicable | `CCT_CIT_TOTAL` | | `classif2.label` | `string` | — | `Citizenship: Total` | | `time` | `int64` | Observation year | `2024` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `108.247` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `—` | | `note_classif.label` | `string` | — | `—` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-age-cct-nb-unemployment-by-sex-age-and-citizenship-thousands") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_CCT_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_AGE_CCT_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_AGE_CCT_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_age_cct_nb_unemployment_by_sex_age_and_citizenship_thousands_2025, title = {Unemployment by sex, age and citizenship (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_CCT_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-age-cct-nb-unemployment-by-sex-age-and-citizenship-thousands}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_AGE_CCT_NB_

This dataset contains 66,202 observations of International migrant stock data across 40 Europe countries, spanning 1991–2025, covering 1 distinct indicator: UNE_TUNE_SEX_AGE_CCT_NB — Unemployment by sex, age and citizenship (thousands). The data is sourced from ILOSTAT, the ILOs central statistics database, pulled via REST API and filtered to Europe ISO3 country codes. It includes columns such as country code, source, indicator, sex, age classification, citizenship classification, time, and observed value, suitable for tabular classification, regression, and time-series forecasting tasks. The data is harmonized using ICLS definitions, with annual frequency, best source selection, and disaggregation dimensions noted for quality assurance.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-age-cct-nb-unemployment-by-sex-age-and-citizenship-thousands 数据集图片
构建方式
该数据集由Electric Sheep Europe团队从国际劳工组织(ILO)的ILOSTAT数据库核心API中直接抽取,仅保留欧洲40个国家的ISO3代码对应数据。原始数据基于各国劳动力调查、家庭收入调查等行政记录,并经ILO统计部门依据国际劳工统计学家会议(ICLS)定义进行标准化处理。抽取的指标为'UNE_TUNE_SEX_AGE_CCT_NB',即按性别、年龄和公民身份分列的失业人数(千计),最终以Parquet格式打包发布,并在HuggingFace上以统一的卡片形式呈现,便于机器学习任务直接调用。
使用方法
用户可通过HuggingFace Datasets库的load_dataset()函数一行加载数据,并直接转换为pandas DataFrame进行后续分析。典型用法包括:按国家代码过滤子集(如df[df['ref_area'] == 'DEU']获取德国数据)、按indicator列筛选后以时间序列形式绘制趋势图、或利用pivot_table构建国家×年份矩阵以进行面板回归或跨国产出比较。数据集列结构清晰,包含ref_area、sex、classif1、classif2、time、obs_value等标准字段,非常适配表格分类、回归及时间序列预测等机器学习任务。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)于2025年通过其ILOSTAT数据库构建,并由Electric Sheep Europe重新封装发布,专注于欧洲40个国家1991至2025年间按性别、年龄和公民身份分层的失业人数(以千计)时序数据。ILOSTAT作为全球劳动统计的权威枢纽,整合了来自劳动力调查、行政记录等多源数据,并依据国际劳工统计学家会议(ICLS)标准进行协调。该数据集共包含66,202条观测记录,其核心研究问题在于揭示欧洲劳动力市场中不同社会人口群体的失业动态差异,为跨国比较、政策评估及劳动经济学研究提供了标准化、细粒度的实证基础。凭借其长达35年的纵向覆盖与多维分类结构,该数据集对分析移民融入、代际就业脆弱性及性别不平等议题具有重要影响力,堪称欧洲区域劳动市场研究的基石资源。
当前挑战
该数据集所解决的领域挑战在于,欧洲劳动力市场因国家间定义差异、调查方法分歧及数据缺口,长期缺乏跨时空可比的失业分层数据。ILOSTAT通过统一标准虽部分化解了此难题,但构建过程中仍面临多重挑战:其一,原始数据源自各国不同调查(如劳动力调查、行政记录),其采样框架与指标口径存在本质性差异,需通过复杂的统计调和处理以消除系统性偏差;其二,时间序列的完整性受限于各国数据采集历史,如部分国家1995年前的记录缺失,且标注了方法修订(如'Break in series'),导致断点衔接与趋势分析困难;其三,分类维度(如年龄组、国籍)的颗粒度在不同国家与年份间未能完全一致,部分观测值因可靠性低被标记为'U'状态,增加了噪声处理与可信度筛选的负担。
常用场景
经典使用场景
该数据集主要用于欧洲劳动力市场研究中失业率的细分分析,尤其关注按性别、年龄和公民身份划分的失业人口数量(单位:千人)。研究者可利用该数据集构建时间序列模型,追踪1991年至2025年间40个欧洲国家的失业动态,或通过面板数据回归探究经济周期、移民政策与劳动力市场结构性变化之间的关联。其多维度分类变量(如性别、年龄组、公民身份)为分解失业率异质性提供了坚实的数据基础,是宏观经济学、劳动经济学及人口迁移研究中不可或缺的量化工具。
解决学术问题
该数据集有效解决了欧洲地区失业数据在时间跨度、国家覆盖及人口细分维度上的碎片化问题,为验证劳动力市场分割理论、移民就业同化假说及性别就业差异提供了统一且可比的标准数据源。其按公民身份划分的失业指标,使学者能够量化外籍劳工与本土劳工在就业冲击中的脆弱性差异,进而评估移民吸纳政策的经济社会影响。此外,结合ILO国际劳工统计标准,该数据推动了跨国面板分析中跨境劳动力流动与失业周期协同性的实证研究。
实际应用
在政策制定与公共管理领域,该数据集被用于设计精准的就业干预措施。欧盟机构及成员国劳动部门可基于性别-年龄-公民身份细分的失业趋势,识别高失业风险群体(如年轻移民女性),从而优化培训补贴与岗位匹配资源分配。国际组织(如ILO、OECD)借助该数据进行跨国基准比较,监测欧洲劳动力市场包容性目标的实现进展。私营部门中,人力资源咨询公司和宏观经济预测机构将其整合进就业景气指数模型,辅助跨国企业制定区域人才招聘与留存策略。
数据集最近研究
最新研究方向
该数据集聚焦于欧洲劳动力市场中分性别、年龄与公民身份的失业率统计,为研究欧洲移民融合、劳动力市场结构性变化以及社会政策制定提供了关键的高精度时间序列数据(1991-2025年,覆盖40国)。近期研究热点集中于利用此类多维分类数据,结合机器学习模型(如时间序列预测与分类算法)分析后疫情时代欧洲各国失业模式的异质性,特别是移民群体与本土公民在就业复苏中的差异化表现。此外,该数据与ILOSTAT标准接轨,为欧盟“技能与人才流动”战略及2030年可持续发展目标(SDG 8:体面工作)的量化监测提供了坚实的实证基础,其开源属性也极大地推动了跨学科合作与政策影响评估的透明化。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务