遇见数据集

electricsheepeurope/europe-ilo-une-deap-sex-geo-dsb-rt-unemployment-rate-by-sex-rural-urban-areas-and-dis

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment rate by sex, rural / urban areas and disability status (%) | Europe (ILOSTAT)" --- # Unemployment rate by sex, rural / urban areas and disability status (%) | Europe (ILOSTAT) 🇪🇺 **13,803 observations** · **30 Europe countries** · **2002–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-13,803-blue) ![countries](https://img.shields.io/badge/countries-30-green) ![years](https://img.shields.io/badge/years-2002–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **13,803 observations** of `Unemployment` data across **30 Europe countries**, spanning **2002–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_GEO_DSB_RT) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_DEAP_SEX_GEO_DSB_RT` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 30 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `ESP` | 566 | 2004 | 2024 | | `AUT` | 565 | 2004 | 2024 | | `FIN` | 560 | 2004 | 2024 | | `FRA` | 560 | 2004 | 2024 | | `PRT` | 557 | 2004 | 2024 | | `ITA` | 556 | 2004 | 2024 | | `GRC` | 553 | 2004 | 2024 | | `LUX` | 541 | 2004 | 2024 | | `SVK` | 540 | 2005 | 2024 | | `HUN` | 536 | 2005 | 2024 | | `EST` | 536 | 2004 | 2024 | | `POL` | 534 | 2005 | 2024 | | `CZE` | 532 | 2005 | 2024 | | `BEL` | 527 | 2004 | 2024 | | `LVA` | 525 | 2005 | 2024 | | ... | _15 more countries_ | | | ## Indicators (sample) - `UNE_DEAP_SEX_GEO_DSB_RT` — Unemployment rate by sex, rural / urban areas and disability status (%) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:7401` | | `source.label` | `string` | Source name in English | `HIES - Living Standards Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_DEAP_SEX_GEO_DSB_RT` | | `indicator.label` | `string` | Indicator name in English | `Unemployment rate by sex, rural / urb…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `GEO_COV_NAT` | | `classif1.label` | `string` | — | `Area type: National` | | `classif2` | `string` | Second classification variable where applicable | `DSB_STATUS_TOTAL` | | `classif2.label` | `string` | — | `Disability status: Total` | | `time` | `int64` | Observation year | `2012` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `22.032` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C14:6260` | | `note_classif.label` | `string` | — | `Nonstandard definition of disability:…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-deap-sex-geo-dsb-rt-unemployment-rate-by-sex-rural-urban-areas-and-dis") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_DEAP_SEX_GEO_DSB_RT"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_DEAP_SEX_GEO_DSB_RT") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_DEAP_SEX_GEO_DSB_RT"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_deap_sex_geo_dsb_rt_unemployment_rate_by_sex_rural_urban_areas_and_dis_2025, title = {Unemployment rate by sex, rural / urban areas and disability status (%) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_GEO_DSB_RT}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-deap-sex-geo-dsb-rt-unemployment-rate-by-sex-rural-urban-areas-and-dis}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_GEO_DSB_RT_

This dataset contains unemployment rate statistics from the International Labour Organization (ILO) ILOSTAT database for 30 European countries spanning 2002 to 2025, with 13,803 observations. It focuses on a single indicator UNE_DEAP_SEX_GEO_DSB_RT, which represents the unemployment rate by sex, rural/urban areas, and disability status (%). The data is sourced directly from the ILOSTAT REST API, filtered to European ISO3 country codes, and harmonized using International Conference of Labour Statisticians (ICLS) definitions. The dataset includes detailed disaggregation dimensions such as sex (total, male, female), area type (e.g., national), and disability status (total), along with columns for data source, observation status, notes, and more. It is suitable for tabular classification, regression, and time-series forecasting tasks. The data is published at an annual frequency, using the ILO-selected best source for consistency and traceability.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-deap-sex-geo-dsb-rt-unemployment-rate-by-sex-rural-urban-areas-and-dis 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT核心统计数据库,基于各成员国劳动力调查、住户收支调查及行政记录等微观数据,经ILO统计部门按照国际劳工统计学家会议定义进行标准化处理与整合。数据通过ILOSTAT REST API直接提取,并过滤出30个欧洲国家的ISO3国家代码,最终形成涵盖2002至2025年间13,803条观测的表格型数据集,每条记录均附有来源标签以确保可追溯性。
特点
数据集的独特之处在于其多维拆解视角,不仅按性别(总、男、女)细分失业率,还同时纳入了城乡地域类型与残疾状态两类分类变量,实现了对劳动力市场结构性差异的精细刻画。此外,每个观测值均附带丰富的元数据列,包括数据来源、观测状态标记(如临时性或不可靠)、以及断点或定义变更等注释信息,为数据质量评估与稳健分析提供了坚实支撑。
使用方法
借助HuggingFace Datasets库,可通过一行代码快速加载数据并转换为Pandas DataFrame进行后续操作。使用者可依据国家代码筛选特定地区的时序数据,或针对单一失业率指标绘制时间序列图以观察趋势变化。数据亦支持按年份与国家列进行透视,生成国家×年份的矩阵表格,便于进行面板数据分析或跨国家比较研究。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)统计司于2025年通过Electric Sheep Europe团队重新整理发布,聚焦欧洲30个国家2002至2025年间按性别、城乡区域及残疾状况分层的失业率数据,共包含13,803条观测记录。作为ILOSTAT数据库的核心子集,它基于国际劳工统计学家会议(ICLS)定义对各国劳动力调查微观数据进行统一归化处理,旨在揭示多重社会维度下劳动力市场的结构性差异。该数据集对劳动经济学、社会政策评估及可持续发展目标(SDG)中的体面工作指标研究具有重要支撑作用,为跨国家、跨时期的系统性比较分析提供了标准化数据基础。
当前挑战
该数据集所解决的领域问题在于量化分析性别、城乡地理与残疾身份三重交叉因素对失业率的叠加影响,弥补传统单一维度失业率统计对弱势群体就业困境的刻画不足。其构建过程中面临多重挑战:首先,30国数据源异构性强,各国劳动力调查的抽样方法、问卷设计及残疾界定标准存在显著差异(例如notes中标记的非标准残疾定义);其次,ILOSTAT虽通过最佳来源选择与ICLS定义归化,但时间序列中仍存在因方法论修订导致的断点(如notes中的系列中断标记),需用户谨慎处理指标连续性;此外,分类变量(性别、区域、残疾状态)的非空约束及部分观测值可靠性标志(如'U'标记为不可靠)也增加了数据清洗与质量控制的复杂性。
常用场景
经典使用场景
在劳动经济学与公共政策研究领域,该数据集被广泛用于分析欧洲各国失业率的时空演变特征。研究者可借助其涵盖2002至2025年、横跨30个欧洲国家的长时序面板数据,构建多维度失业率预测模型,如基于时间序列的季节性分解与自回归移动平均模型,或融合性别、城乡区域及残疾状况等分层变量的多元回归分析。数据集提供的细粒度分类维度,使得深入探究不同人群(如女性、农村居民或残疾群体)在劳动力市场中的脆弱性成为可能。
实际应用
在实际应用层面,该数据集为欧盟委员会及各成员国劳动部门的政策制定提供了数据驱动的决策支持。基于其提供的年度就业率指标,政府机构能够动态监测残疾人就业促进计划的实施成效,并针对农村女性等复合弱势群体优化职业培训资源配置。此外,国际组织如国际劳工组织(ILO)利用此数据定期发布《全球就业与社会展望》报告,用于评估可持续就业目标的区域进展,同时为商业银行与投资机构评估地区经济活力与劳动力成本风险提供量化依据。
衍生相关工作
该数据集衍生了一系列开创性学术工作,尤其在机器学习与因果推断交叉领域表现突出。例如,研究者利用其分层结构开发了面向非平衡面板数据的去偏回归树模型,用于估计残疾歧视对失业持续期的因果效应。此外,融合时空图神经网络与注意力机制的方法被提出,以建模欧洲各国失业率在城乡梯度上的空间溢出效应。相关成果还催生了针对离散时间生存分析的基准测试框架,用于比较不同统计模型在预测边缘化群体再就业概率时的偏差与方差权衡。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务