遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-ocu-nb-unemployment-of-previously-employed-persons-by-sex

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment of previously employed persons by sex and former occupation (thousands) | Europe (ILOSTAT)" --- # Unemployment of previously employed persons by sex and former occupation (thousands) | Europe (ILOSTAT) 🇪🇺 **47,064 observations** · **40 Europe countries** · **1970–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-47,064-blue) ![countries](https://img.shields.io/badge/countries-40-green) ![years](https://img.shields.io/badge/years-1970–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **47,064 observations** of `Unemployment` data across **40 Europe countries**, spanning **1970–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_OCU_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_OCU_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 40 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `AUT` | 1,899 | 1972 | 2025 | | `ESP` | 1,889 | 1976 | 2025 | | `BEL` | 1,846 | 1979 | 2024 | | `DEU` | 1,629 | 1992 | 2024 | | `GRC` | 1,575 | 1981 | 2024 | | `ITA` | 1,570 | 1992 | 2024 | | `CZE` | 1,564 | 1993 | 2024 | | `IRL` | 1,548 | 1992 | 2024 | | `POL` | 1,528 | 1995 | 2025 | | `DNK` | 1,527 | 1992 | 2024 | | `CHE` | 1,516 | 1991 | 2025 | | `RUS` | 1,503 | 1997 | 2025 | | `PRT` | 1,485 | 1974 | 2024 | | `SVK` | 1,481 | 1994 | 2024 | | `HUN` | 1,403 | 1994 | 2024 | | ... | _25 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_OCU_NB` — Unemployment of previously employed persons by sex and former occupation (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:480` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_OCU_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment of previously employed p…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `OCU_SKILL_TOTAL` | | `classif1.label` | `string` | — | `Occupation (Skill level): Total` | | `time` | `int64` | Observation year | `2015` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `91.211` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C4:3052` | | `note_classif.label` | `string` | — | `Nonstandard occupation: Including fir…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `S5:38_T3:104` | | `note_source.label` | `string` | — | `Population coverage: Excluding both i…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-ocu-nb-unemployment-of-previously-employed-persons-by-sex") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_OCU_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_OCU_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_OCU_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_ocu_nb_unemployment_of_previously_employed_persons_by_sex_2025, title = {Unemployment of previously employed persons by sex and former occupation (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_OCU_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-ocu-nb-unemployment-of-previously-employed-persons-by-sex}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_OCU_NB_

This dataset contains unemployment data from the International Labour Organization (ILO) ILOSTAT database, specifically the indicator Unemployment of previously employed persons by sex and former occupation (thousands). It covers 40 European countries from 1970 to 2025, with 47,064 observations. Data is sourced via the ILOSTAT REST API and filtered to Europe ISO3 country codes, harmonized using International Conference of Labour Statisticians (ICLS) definitions. The dataset includes structured columns such as country codes, indicator codes, sex disaggregation, year, observed values, and status flags, suitable for tabular classification, regression, and time-series forecasting tasks. Data is published at annual frequency and includes quality caveats, such as the use of ILO-selected best source for multiple sources.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-ocu-nb-unemployment-of-previously-employed-persons-by-sex 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT统计数据库,专注于欧洲地区先前就业人员按性别与先前职业划分的失业人数(单位:千)。数据通过ILOSTAT REST API接口直接拉取,原始指标代码为UNE_TUNE_SEX_OCU_NB,并依据欧洲ISO3国家代码进行地理范围过滤。ILOSTAT对各国的劳动力调查微观数据进行了标准化处理,遵循国际劳工统计学家会议(ICLS)的定义,同时通过source.label字段标注数据来源,确保数据可追溯。Electric Sheep Europe团队完成了数据的重新封装与整合,生成了包含47,064条观测值、覆盖40个欧洲国家、时间跨度为1970年至2025年的结构化数据集。
特点
本数据集最显著的特点在于其精细的维度划分与广泛的地理覆盖。除基础的失业数值外,数据按性别(总、男性、女性)及先前职业(基于技能水平分类)进行多维度分解,为分析不同人口群体的失业结构提供了宝贵素材。数据集涵盖了40个欧洲国家长达55年的年度观测,时间序列完整,且每条记录均包含观测状态标志(如临时性或不可靠数据)及详尽的数据注释,有助于研究者评估数据质量与潜在的方法论断裂。所有数据均遵循CC-BY-4.0许可协议,便于学术与商业应用。
使用方法
研究者可通过HuggingFace Datasets库的load_dataset函数快速加载数据,并将其转换为Pandas DataFrame进行后续分析。典型的使用场景包括:筛选特定国家(如德国DEU)的失业时间序列,按性别或职业分组考察失业率动态;也可利用pivot_table将数据重塑为国家×年份的矩阵形式,便于进行面板数据回归或跨国家比较。由于数据以年度频率发布,建议在分析时注意季节性与短期波动。同时,用户应参考obs_status和note_indicator字段,对标注为不可靠或存在方法修订的观测值进行适当过滤或稳健性检验。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)于2025年通过其核心统计数据库ILOSTAT创建,并由Electric Sheep Europe团队重新整理与封装,聚焦于1970年至2025年间欧洲40个国家中此前有就业经历人群的失业状况,依据性别与先前职业进行细分,共包含47,064条观测记录。在劳动经济学与公共政策研究领域,失业问题始终是衡量劳动力市场健康与否的关键指标,而基于性别与职业的精细分层数据,为揭示结构性失业、性别就业差异以及职业流动性等深层议题提供了不可替代的数据基础。ILO作为全球劳动统计的权威机构,其数据经过严格的方法论协调与跨国可比性处理,使该数据集成为研究欧洲劳动力市场演变、评估就业政策效果以及构建预测模型的宝贵资源,对推动基于证据的社会科学研究具有重要影响力。
当前挑战
该数据集所应对的核心领域挑战在于:其一,如何准确量化并捕捉欧洲各国在长达半个多世纪中,因经济周期、产业转型与政策变革导致的失业形态变化,尤其是性别与职业维度的异质性影响;其二,数据构建过程中面临来源多元、定义不一的困境——各国劳动力调查方法、职业分类标准及统计口径存在差异,ILOSTAT通过国际劳工统计学家会议(ICLS)定义进行协调,但数据质量标注(如“不可靠”、“方法修订”)提示了跨时空比较的内在复杂性。此外,部分指标仅以年度频率呈现,缺少更细粒度的时间分解,且当同一国家与年份存在多个数据源时,ILO选用的“最佳来源”策略虽确保了权威性,却也可能隐含着不可忽视的统计偏差,这些均对研究者的数据甄别与因果推断能力提出了严峻考验。
常用场景
经典使用场景
该数据集汇聚了1970年至2025年间40个欧洲国家按性别和先前职业分类的失业人数观测值,共计47,064条记录,源自国际劳工组织ILOSTAT数据库。其最经典的用途在于构建时间序列模型,用以分析欧洲劳动力市场中不同性别与职业群体的失业动态演变趋势。研究者可借助此数据揭示经济周期、产业结构变迁以及社会政策调整对特定人群就业状况的差异化影响,从而为劳动经济学中的结构性失业与摩擦性失业研究提供坚实的数据支撑。
衍生相关工作
围绕该数据集,衍生出了一系列深入的学术与实践工作。在模型层面,研究者基于此数据开发了多变量时间序列预测模型,如结合季节性分解与长短时记忆网络的混合模型,以提升短期失业率预测精度。在理论层面,衍生了探讨欧洲一体化进程中劳动力市场趋同或分化的实证研究,以及评估新冠疫情对不同性别和职业群体失业冲击的非对称性影响。此外,该数据集还促进了跨数据库的融合分析,例如与薪资数据、职业培训投入数据结合,构建起劳动力市场供需匹配效率的综合评估框架。
数据集最近研究
最新研究方向
该数据集聚焦于欧洲地区先前就业者失业状况的性别与职业细分,为劳动经济学与时间序列预测领域提供了高质量的结构化数据支撑。当前研究前沿围绕后疫情时代欧洲劳动力市场的结构性转型,尤其是性别差异与职业技能错配对失业动态的交互影响。伴随着欧盟“绿色新政”与数字化转型加速,该数据可助力探究传统制造业与新兴服务业的就业更迭脉络,为跨国政策模拟与预测模型提供基准。其长达五十余年的跨度也使之成为评估长周期社会事件(如金融危机、疫情冲击)对就业韧性影响的宝贵实证资源。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务