遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-ocu-edu-nb-unemployment-by-sex-occupation-and-education-thous

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment by sex, occupation and education (thousands) | Europe (ILOSTAT)" --- # Unemployment by sex, occupation and education (thousands) | Europe (ILOSTAT) 🇪🇺 **54,829 observations** · **37 Europe countries** · **1991–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-54,829-blue) ![countries](https://img.shields.io/badge/countries-37-green) ![years](https://img.shields.io/badge/years-1991–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **54,829 observations** of `Unemployment` data across **37 Europe countries**, spanning **1991–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_OCU_EDU_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_OCU_EDU_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 37 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GBR` | 2,279 | 1992 | 2020 | | `IRL` | 2,149 | 1992 | 2024 | | `DEU` | 2,120 | 1992 | 2024 | | `GRC` | 2,110 | 1993 | 2020 | | `ESP` | 2,100 | 1992 | 2025 | | `DNK` | 2,095 | 1992 | 2024 | | `ITA` | 2,094 | 1992 | 2024 | | `BEL` | 2,023 | 1992 | 2024 | | `PRT` | 1,933 | 1992 | 2024 | | `CHE` | 1,822 | 1991 | 2025 | | `AUT` | 1,703 | 1995 | 2025 | | `POL` | 1,698 | 1997 | 2025 | | `CZE` | 1,696 | 1998 | 2024 | | `HUN` | 1,555 | 1997 | 2024 | | `SVN` | 1,547 | 1996 | 2024 | | ... | _22 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_OCU_EDU_NB` — Unemployment by sex, occupation and education (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:480` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_OCU_EDU_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, occupation and e…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `OCU_SKILL_TOTAL` | | `classif1.label` | `string` | — | `Occupation (Skill level): Total` | | `classif2` | `string` | Second classification variable where applicable | `EDU_AGGREGATE_TOTAL` | | `classif2.label` | `string` | — | `Education (Aggregate levels): Total` | | `time` | `int64` | Observation year | `2010` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `77.239` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:5578` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-ocu-edu-nb-unemployment-by-sex-occupation-and-education-thous") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_OCU_EDU_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_OCU_EDU_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_OCU_EDU_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_ocu_edu_nb_unemployment_by_sex_occupation_and_education_thous_2025, title = {Unemployment by sex, occupation and education (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_OCU_EDU_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-ocu-edu-nb-unemployment-by-sex-occupation-and-education-thous}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_OCU_EDU_NB_

This dataset contains unemployment statistics from the ILOSTAT database of the International Labour Organization (ILO), specifically the indicator Unemployment by sex, occupation and education (thousands). It covers 37 European countries from 1991 to 2025, with 54,829 observations. Data is sourced via the ILOSTAT REST API and harmonized using International Conference of Labour Statisticians (ICLS) definitions. The dataset includes dimensions such as country, year, sex, occupation classification, and education level, along with source information, observation status, and notes. Organized in tabular format, it is suitable for tasks like tabular classification, regression, and time-series forecasting. Repackaged by Electric Sheep Europe for machine learning readiness, it is released under the CC-BY-4.0 license.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-ocu-edu-nb-unemployment-by-sex-occupation-and-education-thous 数据集图片
构建方式
该数据集基于国际劳工组织(ILO)的ILOSTAT REST API直接抽取,原始数据源自各国劳动力调查、家庭收入调查、机构调查及行政记录等多元渠道,经ILO统计部门依据国际劳工统计学家会议(ICLS)定义进行标准化与整合。数据采集后,按照欧洲ISO3国家代码进行地理过滤,最终汇聚成涵盖37个欧洲国家、时间跨度从1991年至2025年的失业率观测值,共计54,829条记录。数据集中包含性别、职业技能等级、教育水平等关键分类维度,并附有数据来源标签以保障可追溯性。
使用方法
该数据集通过HuggingFace Datasets库轻松加载,用户可使用`load_dataset`函数直接获取数据,并转换为Pandas DataFrame进行后续分析。典型应用包括按国家筛选数据以聚焦特定区域,或通过排序与绘图对特定指标进行时间序列可视化。研究者还能借助数据透视表操作,构建以时间为行、国家为列的矩阵格式,便于进行跨国家比较和面板数据计量分析。数据以Parquet格式打包,兼顾存储效率与读取速度,适合机器学习与统计分析的工作流。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)下属的ILOSTAT统计数据库于2025年发布,经Electric Sheep Europe团队重新封装并托管于HuggingFace平台,旨在提供欧洲37个国家1991至2025年间按性别、职业和教育程度划分的失业人口数(单位:千人)。核心研究问题聚焦于劳动力市场结构性差异的量化分析,为跨国家、跨时期的失业动态比较提供标准化数据基础。作为全球劳动统计领域的权威来源,ILOSTAT整合了各国劳动力调查、行政记录等多源微观数据,其发布的失业指标在劳动经济学、公共政策评估及可持续发展目标监测中具有广泛影响力,该数据集进一步通过统一格式与便捷接口,显著降低了研究者获取高质量面板数据的技术门槛。
当前挑战
该数据集解决的领域挑战在于,劳动力市场结构性失业分析长期受限于碎片化数据源与不一致的分类标准,尤其是性别、职业技能水平与教育层级的多维交叉细分数据难以获取,阻碍了精准政策干预的制定。构建过程中面临的挑战包括:需协调37国不同劳动统计体系下的调查方法与定义差异,ILO依靠ICLS国际标准进行数据协调与最佳来源筛选,但数据标注仍需通过'obs_status'字段标记'不可靠'观测值;时间跨度长达35年,期间各国方法论修订导致序列非连续性,需借助'note_indicator'标注断点;此外,部分维度的缺失值问题(如分类列仅在对应指标发布时有效)要求用户谨慎处理分层聚合,以保证时间序列建模的稳健性。
常用场景
经典使用场景
在劳动经济学与公共政策研究领域,该数据集最经典的用途是支撑跨国失业动态的深度剖析。研究者可借助其涵盖37个欧洲国家、横跨1991至2025年的长时序面板数据,精准刻画不同性别、职业与教育水平群体的失业率演变轨迹。通过按性别(总、男、女)及职业技能等级与教育程度进行交叉分组,学者能够构建多维度的劳动力市场状态模型,从而系统探究欧洲各国结构性失业的异质性根源,为比较制度分析提供扎实的数据基石。
解决学术问题
该数据集的核心学术贡献在于破解了传统失业统计中维度单一、国际可比性不足的困境。它借助国际劳工组织(ILO)统一协调的ICLS定义与标准化分类体系,使得跨国家、跨时期的性别-职业-教育三维失业结构得以严谨对照。研究者可据此检验人力资本理论在失业风险分配中的普适性,辨识技能错配与教育回报率波动对就业市场的长期冲击,同时揭示福利制度与劳动力市场管制如何调节不同群体的失业脆弱性,从而推动劳动经济学的实证边界向精细化与动态化方向拓展。
实际应用
在实际应用层面,该数据集为欧盟及各成员国的就业政策制定与效果评估提供了量化支撑。政策分析师可以基于分性别、分职业和分教育水平的失业统计,精准定位受技能转型冲击最为严重的群体,从而设计具有针对性的再培训计划与职业引导策略。此外,劳动力市场监测机构能够利用该时序数据构建预警模型,实时捕捉地区性或行业性失业波动的前兆信号,为宏观经济调控与社会保障资源的优化配置提供科学依据,有效提升公共治理的响应效率。
数据集最近研究
最新研究方向
基于欧洲37国1991至2025年间按性别、职业与教育层次细分的失业数据,该数据集为劳动力市场结构化失衡与技能错配的时序分析提供了宝贵的颗粒化语料。近期前沿研究聚焦于利用此类高维分类面板数据,结合深度学习模型(如Transformer或时间序列聚类)来推演后疫情时代欧洲各国失业率的异质性演化路径,并量化自动化和绿色转型对不同性别与教育群体就业冲击的差异化效应。该数据与欧盟“2030数字十年”及“技能公约”等政策议程紧密关联,为实证评估教育培训干预对缓解结构性失业的效能提供了基准支持。其长达34年的跨周期覆盖和ILOSTAT的严格同质化处理,对于构建稳健的劳动力市场预警系统与推动因果推断方法在劳动经济学中的应用具有显著的方法论意义。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务