遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-geo-nb-unemployment-by-sex-and-rural-urban-areas-thousand

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 1K<n<10K tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment by sex and rural / urban areas (thousands) | Europe (ILOSTAT)" --- # Unemployment by sex and rural / urban areas (thousands) | Europe (ILOSTAT) 🇪🇺 **8,188 observations** · **39 Europe countries** · **1987–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-8,188-blue) ![countries](https://img.shields.io/badge/countries-39-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **8,188 observations** of `Unemployment` data across **39 Europe countries**, spanning **1987–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_GEO_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_GEO_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 39 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GRC` | 351 | 1987 | 2025 | | `FRA` | 297 | 1992 | 2024 | | `BEL` | 297 | 1992 | 2024 | | `DNK` | 297 | 1992 | 2024 | | `ESP` | 297 | 1992 | 2024 | | `DEU` | 297 | 1992 | 2024 | | `ITA` | 297 | 1992 | 2024 | | `PRT` | 297 | 1992 | 2024 | | `NLD` | 297 | 1992 | 2024 | | `IRL` | 288 | 1993 | 2024 | | `AUT` | 279 | 1995 | 2025 | | `FIN` | 276 | 1995 | 2024 | | `GBR` | 274 | 1992 | 2019 | | `SWE` | 270 | 1995 | 2024 | | `LUX` | 252 | 1997 | 2024 | | ... | _24 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_GEO_NB` — Unemployment by sex and rural / urban areas (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:7401` | | `source.label` | `string` | Source name in English | `HIES - Living Standards Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_GEO_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex and rural / urban…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `GEO_COV_NAT` | | `classif1.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2012` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `207.786` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C12:2951` | | `note_classif.label` | `string` | — | `Rural / urban areas coverage definiti…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-geo-nb-unemployment-by-sex-and-rural-urban-areas-thousand") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_GEO_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_GEO_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_GEO_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_geo_nb_unemployment_by_sex_and_rural_urban_areas_thousand_2025, title = {Unemployment by sex and rural / urban areas (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_GEO_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-geo-nb-unemployment-by-sex-and-rural-urban-areas-thousand}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_GEO_NB_

This dataset contains 8,188 observations of unemployment data across 39 Europe countries, spanning from 1987 to 2025, covering one distinct indicator: Unemployment by sex and rural / urban areas (thousands). The data is sourced from the International Labour Organization (ILO) ILOSTAT database, retrieved via the ILOSTAT REST API and filtered to Europe country codes. It includes columns such as country code, country name, data source, indicator code, sex disaggregation, year, observed value, and provides data quality notes and usage examples, suitable for tabular classification, regression, and time-series forecasting tasks.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-geo-nb-unemployment-by-sex-and-rural-urban-areas-thousand 数据集图片
构建方式
该数据集的构建依托于国际劳工组织(ILO)的ILOSTAT统计数据库,通过REST API直接获取原始指标数据。数据源为官方指标体系中的`UNE_TUNE_SEX_GEO_NB`指标,代表按性别和城乡区域划分的失业人数(单位:千)。原始数据经由ILO依据国际劳工统计学家会议(ICLS)定义进行标准化处理,并标记了数据来源以保障可追溯性。Electric Sheep Europe团队进一步对数据进行过滤,仅保留欧洲39个国家的观测值,最终整合为包含8,188条记录的时间序列数据集,覆盖1987年至2025年的时间跨度。
特点
该数据集的核心特点在于其多维度结构设计与精细化粒度。数据不仅提供了按性别(总、男、女)和城乡区域(国家层面)划分的失业统计,还保留了丰富的元数据列,如来源标识、观测状态标志以及各类注释(如方法修订、数据覆盖定义),为研究者在分析时进行质量评估提供了关键信息。作为面向欧洲区域的专门化子集,它整合了39个国家的数据,具有较好的地理代表性,同时兼顾了时间序列的长期性,非常适合进行跨国的劳动市场对比与宏观趋势建模。
使用方法
数据集的使用极为便捷,研究者可通过Hugging Face的`datasets`库直接加载。加载后的数据可转换为Pandas DataFrame进行操控,例如通过`ref_area`列按国家筛选,如筛选德国数据;也可基于`time`和`obs_value`列绘制特定指标的时序图。高级分析场景中,可采用透视表操作,将数据重塑为以年份为行、国家为列的矩阵形式,便于进行面板数据分析或多国横向对比。该数据集结构规整,适配表格分类、回归及时间序列预测等多种机器学习任务。
背景与挑战
背景概述
失业率作为衡量劳动力市场健康程度的关键指标,其空间异质性分析对于制定精准的区域就业政策至关重要。由国际劳工组织(ILO)统计司创建并维护的ILOSTAT数据库,作为全球劳动统计领域的权威数据源,于2025年经Electric Sheep Europe团队重新封装后发布此数据集。该数据集聚焦欧洲39个国家,涵盖了1987至2025年间按性别与城乡地域细分的失业人数(以千计),共包含8,188条观测记录。其核心研究问题在于揭示欧洲不同国家、性别及城乡维度下失业模式的时空演变规律,为劳动经济学、区域发展与公共政策研究提供了标准化的高价值数据基础,对推动欧洲劳动力市场的跨时空比较分析具有显著影响力。
当前挑战
该数据集所解决的领域挑战在于,传统的失业统计往往仅提供全国总体数据,而忽略了性别与城乡地域的异质性,导致政策干预缺乏针对性。ILOSTAT通过整合各国劳动力调查、家庭收支调查等多元来源,并依据国际劳工统计学家会议(ICLS)定义进行标准化,有效缓解了跨国数据可比性不足的问题。然而,数据构建过程面临严峻挑战:各国统计能力参差不齐,数据来源标签(source.label)虽标注了来源类型,但不同调查方法间的系统性偏差难以完全消除;时间序列的连续性受制于各国统计体系修订(如note_indicator.label中提及的‘Break in series: Methodology revised’)导致的断点问题;此外,部分观测值因数据质量被标记为‘不可靠’,增加了模型训练时噪声处理的复杂性。
常用场景
经典使用场景
该数据集为欧洲39个国家1987至2025年间按性别和城乡划分的失业人数提供了高质量的年度观测数据,共计8188条记录。在学术研究中,它常被用于构建面板数据模型,以分析欧洲各国劳动力市场的结构性差异与动态变化。经典使用包括利用时间序列预测技术对特定国家或区域的失业趋势进行中长期推演,或基于分类与回归方法探究性别、城乡地域等维度对失业率的影响权重。数据经过国际劳工组织(ILO)依据ICLS定义进行标准化处理,来源清晰且标注详尽,使其成为跨国家、跨时段比较研究的坚实基石,尤其适用于揭示欧洲内部失业模式的空间异质性与时间演化规律。
衍生相关工作
基于该数据集,一系列衍生性学术工作得以展开。研究者结合性别与城乡双重维度,构建了融合空间计量经济学与时间序列分析的综合框架,用以剖析失业率的区域溢出效应与跨群体传导机制。部分工作将ILOSTAT的这一欧洲子集与微观调查数据、工资面板或产业就业数据相链接,形成多源融合的劳动市场分析数据库,从而支持对失业成因与影响因素的更立体建模。此外,该数据集也被用作基准测试平台,用于评估新提出的时变系数面板模型或非平稳面板协整方法的表现,尤其是在存在结构断点与国家异质性的情境下,推动了计量方法论的前沿探索。
数据集最近研究
最新研究方向
该数据集聚焦于欧洲39国1987至2025年间按性别及城乡区域划分的失业率动态,为劳动经济学与时空数据分析提供了精细化颗粒度的支撑。在区域经济不平等与绿色转型交织的当下,研究者正利用此类高分辨率面板数据,结合机器学习与时间序列模型,剖析就业结构对政策冲击与劳动力迁移的响应。其多维度分类特征(性别、城乡)尤其契合了国际劳工组织与欧盟统计局对包容性增长与代际公平的量化诉求,成为探索后疫情时代欧洲就业韧性、技术替代效应及农村振兴路径的关键基础资源,助力揭示隐含于宏观统计之下的结构性失衡与演化规律。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务