遇见数据集

electricsheepeurope/europe-ilo-une-deap-sex-geo-mts-rt-unemployment-rate-by-sex-rural-urban-area-and-mari

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment rate by sex, rural / urban area and marital status (%) | Europe (ILOSTAT)" --- # Unemployment rate by sex, rural / urban area and marital status (%) | Europe (ILOSTAT) 🇪🇺 **28,602 observations** · **37 Europe countries** · **1987–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-28,602-blue) ![countries](https://img.shields.io/badge/countries-37-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **28,602 observations** of `Unemployment` data across **37 Europe countries**, spanning **1987–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_GEO_MTS_RT) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_DEAP_SEX_GEO_MTS_RT` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 37 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `POL` | 1,869 | 2000 | 2025 | | `MDA` | 1,816 | 2000 | 2025 | | `AUT` | 1,794 | 1995 | 2025 | | `FRA` | 1,546 | 2005 | 2024 | | `CHE` | 1,363 | 2000 | 2025 | | `BIH` | 1,350 | 2001 | 2020 | | `RUS` | 1,296 | 2010 | 2025 | | `SRB` | 1,013 | 2007 | 2020 | | `GRC` | 918 | 1987 | 2020 | | `GBR` | 823 | 1992 | 2019 | | `BLR` | 803 | 2016 | 2024 | | `BEL` | 783 | 1992 | 2020 | | `ESP` | 783 | 1992 | 2020 | | `ITA` | 783 | 1992 | 2020 | | `DNK` | 783 | 1992 | 2020 | | ... | _22 more countries_ | | | ## Indicators (sample) - `UNE_DEAP_SEX_GEO_MTS_RT` — Unemployment rate by sex, rural / urban area and marital status (%) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:7401` | | `source.label` | `string` | Source name in English | `HIES - Living Standards Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_DEAP_SEX_GEO_MTS_RT` | | `indicator.label` | `string` | Indicator name in English | `Unemployment rate by sex, rural / urb…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `GEO_COV_NAT` | | `classif1.label` | `string` | — | `Area type: National` | | `classif2` | `string` | Second classification variable where applicable | `MTS_AGGREGATE_TOTAL` | | `classif2.label` | `string` | — | `Marital status (Aggregate): Total` | | `time` | `int64` | Observation year | `2012` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `22.032` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-deap-sex-geo-mts-rt-unemployment-rate-by-sex-rural-urban-area-and-mari") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_DEAP_SEX_GEO_MTS_RT"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_DEAP_SEX_GEO_MTS_RT") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_DEAP_SEX_GEO_MTS_RT"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_deap_sex_geo_mts_rt_unemployment_rate_by_sex_rural_urban_area_and_mari_2025, title = {Unemployment rate by sex, rural / urban area and marital status (%) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_GEO_MTS_RT}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-deap-sex-geo-mts-rt-unemployment-rate-by-sex-rural-urban-area-and-mari}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_DEAP_SEX_GEO_MTS_RT_

This dataset contains unemployment rate data from the International Labour Organization (ILO) ILOSTAT database for Europe, specifically the indicator Unemployment rate by sex, rural / urban area and marital status (%). It covers 37 European countries from 1987 to 2025, with 28,602 observations. The data is provided in tabular format, including columns such as country code, year, observed value, data source, sex disaggregation, rural/urban area classification, marital status classification, and data quality notes (e.g., annual frequency, best source selection, non-null disaggregation columns). The dataset is repackaged by Electric Sheep Europe as part of a unified, ML-ready data layer for Europe.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-deap-sex-geo-mts-rt-unemployment-rate-by-sex-rural-urban-area-and-mari 数据集图片
构建方式
本数据集源自国际劳工组织(ILO)的ILOSTAT统计数据库,通过REST API接口直接抽取原始指标数据,并依据欧洲ISO3国家代码进行地理范围过滤,最终整合为涵盖37个欧洲国家、时间跨度从1987年至2025年的28,602条观测记录。数据构建过程中严格遵循国际劳工统计学家会议(ICLS)的定义标准,对原始调查微观数据进行协调统一,并在`source.label`列中明确标注数据来源,确保每一条记录均可追溯。
特点
该数据集的核心特色在于其精细化的多维度分层结构,不仅按性别(总、男、女)对失业率进行分解,还引入了`classif1`和`classif2`分类变量,分别标识地理覆盖范围(如全国、城乡区域)和婚姻状况(如总、已婚、单身等),从而能够支持对欧洲劳动力市场进行交叉维度的深入分析。此外,数据还包含了观测状态标志、方法论修订说明及源数据注释等元信息,为用户提供了高透明度的数据质量评估依据。
使用方法
用户可通过HuggingFace Datasets库轻松调用本数据集,使用`load_dataset()`函数即可在数秒内将数据加载为Pandas DataFrame,便于后续的统计分析与建模。典型的使用场景包括按国家筛选特定区域的失业时间序列、针对单一指标进行趋势可视化,或利用数据透视表构建国家-年份矩阵,以实现跨国的面板数据比较。数据集同时适用于表格分类、回归分析及时间序列预测等机器学习任务。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)于2025年通过其ILOSTAT数据库整合发布,并由Electric Sheep Europe团队重新封装,收录了1987年至2025年间欧洲37个国家的28,602条失业率观测记录。数据按性别、城乡区域及婚姻状况进行多维细分,旨在揭示欧洲劳动力市场中不同社会群体间的失业率差异。作为全球劳动统计的核心来源,ILOSTAT通过统一协调各国劳动力调查、家庭收支调查及行政记录数据,为政策制定者与研究者提供了高质量的跨时空劳动力市场指标。该数据集的出现填补了欧洲失业率精细化对比分析的空白,为评估性别平等、城乡发展不均衡及婚姻状况对就业影响等社会经济议题奠定了数据基础,对劳动经济学、人口学及公共政策研究具有重要参考价值。
当前挑战
领域层面,失业率因性别、地域和婚姻状况的交互作用呈现复杂态势,传统宏观统计数据难以揭示子群体间的结构性差异,例如农村女性与城市已婚男性的失业率分化,这需要支持多维度交叉分析的高粒度数据集。构建过程中,主要挑战来自数据源异构性与标准化:不同国家采用各异的劳动力调查框架(如季节性调整、问卷设计差异),ILO虽依据国际劳工统计学家会议(ICLS)定义进行协调,但数据仍存在年度频率不均、观测状态标记不可靠(如临时或估算值)以及因方法论修订导致的序列断点问题。此外,部分国家覆盖年份不连续(如俄罗斯仅自2010年起有数据),且城乡划分标准因国而异,进一步增加了跨域可比性与时间序列分析的难度。
常用场景
经典使用场景
在欧洲劳动经济学与社会政策研究领域,该数据集以ILOSTAT官方统计为基础,整合了37个欧洲国家1987至2025年间按性别、城乡区域及婚姻状况分层的失业率观测值,共计28,602条记录。经典使用场景聚焦于构建面板数据模型,系统分析欧洲劳动力市场中结构性失业的时空演变规律。研究者通常利用其丰富的分类维度,探究性别差异、城乡分割与婚姻状况对失业风险的交互效应,或借助时间序列分解方法挖掘周期性波动与长期趋势。此外,该数据集常被用于训练机器学习分类与回归模型,预测不同人口亚群的失业概率,以及在时序预测任务中评估社会经济指标的动态演化,从而为跨国比较研究提供标准化、机器可读的坚实数据基础。
实际应用
在实际应用层面,该数据集为欧洲各国政府、国际组织及研究机构提供了可操作的决策支持工具。政策制定者可依据性别与城乡细分数据,精准识别失业高风险群体,设计差异化的职业培训与社保干预方案。国际劳工组织与欧盟委员会借助此类标准化的时序指标,监控成员国在充分就业及性别平等方面的进展,评估《欧洲就业战略》的实施成效。金融机构与投资分析公司则将其嵌入宏观经济预测模型,量化失业率对消费、信贷风险与区域经济韧性的影响,优化资产配置策略。数据科学家亦将其用于构建实时劳动力市场预警系统,通过特征工程与时间序列异常检测,为经济景气监测提供前瞻性信号,助力敏捷的政策响应。
衍生相关工作
该数据集已衍生出一系列具有学术影响力的经典工作。在方法论层面,研究者基于其面板结构发展了空间计量与多水平随机前沿模型,精细刻画跨国失业率的空间溢出效应。另有学者将之与ILOSTAT其他数据库(如工资、工时指标)融合,构建多维劳动脆弱性指数,推动劳动力市场综合评估体系的完善。在实证研究集群中,围绕欧洲女性和农村劳动力市场展开的系列论文,借助数据集的精细分维度,重新审视了婚姻回报与性别工资差距在失业语境下的表现。此外,该数据集被用于机器学习领域的可解释性研究,通过树模型与神经网络的注意力机制剖析失业预测的关键驱动力,成为劳动经济学与数据科学交叉创新的典型案例,催生了大量关于社会福利损失与失业持久性的高阶计量研究。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务