遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-geo-dsb-nb-unemployment-by-sex-rural-urban-areas-and-disabili

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment by sex, rural / urban areas and disability status (thousands) | Europe (ILOSTAT)" --- # Unemployment by sex, rural / urban areas and disability status (thousands) | Europe (ILOSTAT) 🇪🇺 **13,803 observations** · **30 Europe countries** · **2002–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-13,803-blue) ![countries](https://img.shields.io/badge/countries-30-green) ![years](https://img.shields.io/badge/years-2002–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **13,803 observations** of `Unemployment` data across **30 Europe countries**, spanning **2002–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_GEO_DSB_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_GEO_DSB_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 30 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `ESP` | 566 | 2004 | 2024 | | `AUT` | 565 | 2004 | 2024 | | `FIN` | 560 | 2004 | 2024 | | `FRA` | 560 | 2004 | 2024 | | `PRT` | 557 | 2004 | 2024 | | `ITA` | 556 | 2004 | 2024 | | `GRC` | 553 | 2004 | 2024 | | `LUX` | 541 | 2004 | 2024 | | `SVK` | 540 | 2005 | 2024 | | `HUN` | 536 | 2005 | 2024 | | `EST` | 536 | 2004 | 2024 | | `POL` | 534 | 2005 | 2024 | | `CZE` | 532 | 2005 | 2024 | | `BEL` | 527 | 2004 | 2024 | | `LVA` | 525 | 2005 | 2024 | | ... | _15 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_GEO_DSB_NB` — Unemployment by sex, rural / urban areas and disability status (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:7401` | | `source.label` | `string` | Source name in English | `HIES - Living Standards Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_GEO_DSB_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment by sex, rural / urban ar…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `GEO_COV_NAT` | | `classif1.label` | `string` | — | `Area type: National` | | `classif2` | `string` | Second classification variable where applicable | `DSB_STATUS_TOTAL` | | `classif2.label` | `string` | — | `Disability status: Total` | | `time` | `int64` | Observation year | `2012` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `207.786` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C14:6260` | | `note_classif.label` | `string` | — | `Nonstandard definition of disability:…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-geo-dsb-nb-unemployment-by-sex-rural-urban-areas-and-disabili") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_GEO_DSB_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_GEO_DSB_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_GEO_DSB_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_geo_dsb_nb_unemployment_by_sex_rural_urban_areas_and_disabili_2025, title = {Unemployment by sex, rural / urban areas and disability status (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_GEO_DSB_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-geo-dsb-nb-unemployment-by-sex-rural-urban-areas-and-disabili}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_GEO_DSB_NB_

This dataset is named "Number of Unemployed Persons (in thousands) classified by Sex, Urban/Rural Area and Disability Status | Europe (ILOSTAT)". It is a tabular dataset focused on unemployment statistics for the European region. The dataset contains 13,803 observations, covering 30 European countries, with a time span from 2002 to 2025. The core indicator is UNE_TUNE_SEX_GEO_DSB_NB, which represents the number of unemployed persons (in thousands) classified by sex, urban/rural area and disability status. The data is sourced from the ILOSTAT database of the International Labour Organization (ILO), obtained via API and filtered to European countries. The dataset provides multi-dimensional categorical variables including sex (total, male, female), region type (e.g., national total), disability status (total), etc., and also includes columns such as data source, observation status and notes. It is suitable for natural language processing tasks such as tabular classification, regression and time series forecasting. The dataset is released in Parquet format, licensed under CC-BY-4.0, and repackaged by Electric Sheep Europe, aiming to provide a unified, machine learning-ready data layer for Europe.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-geo-dsb-nb-unemployment-by-sex-rural-urban-areas-and-disabili 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT核心统计数据库,经由Electric Sheep Europe团队重新封装而成。数据通过ILOSTAT REST API接口直接抽取,具体指标代码为UNE_TUNE_SEX_GEO_DSB_NB,并依据欧洲ISO3国家代码进行地理范围过滤,最终整合了来自30个欧洲国家的13,803条观测记录,时间跨度覆盖2002年至2025年。ILOSTAT原始数据基于各国劳动力调查、家庭收入调查等微观数据,并依据国际劳工统计学家会议(ICLS)定义进行统一协调,数据来源在source.label字段中予以标注,确保了数据的溯源性与可比性。
使用方法
该数据集已整合至HuggingFace Datasets库,用户可通过一行命令`load_dataset()`便捷加载,并直接转换为Pandas DataFrame进行后续分析。典型的使用场景包括:按国家代码过滤以获取单一国家的失业序列,按时间排序绘制特定指标的时序图以观察趋势,以及通过数据透视将原始长格式数据重塑为国家×年份的矩阵,便于进行面板数据分析或比较研究。数据集采用CC-BY-4.0许可协议,用户在使用时需同时引用原始ILO数据来源与Electric Sheep Europe的重新封装版本,以保障知识产权的合规性。
背景与挑战
背景概述
该数据集由国际劳工组织(ILO)于2025年通过其核心统计数据库ILOSTAT创建,并经由Electric Sheep Europe团队重新封装,聚焦于欧洲30个国家2002至2025年间按性别、城乡区域及残疾状况划分的失业人数(单位:千)。作为劳动统计领域的权威来源,ILOSTAT基于各国劳动力调查、家庭收支调查及行政记录数据,经国际劳工统计学家会议(ICLS)定义统一协调后生成。该数据集为研究欧洲劳动力市场中多重边缘化群体的结构性失业问题提供了高粒度、跨时期的面板数据,尤其对评估残疾人群体的就业包容性政策、城乡发展差异及性别不平等具有关键支撑作用,推动了劳动经济学、社会政策及可持续发展目标(SDG体面工作指标)的实证分析。
当前挑战
该数据集面临的挑战首先在于所解决的领域问题:失业率统计常忽视残疾、性别与城乡维度的交叉影响,导致政策干预难以精准定位弱势群体,而该数据通过细化分层揭示了隐蔽的就业歧视。构建过程中的挑战则包括:ILOSTAT需整合各国差异化的调查方法论(如残疾定义的非标准化),导致部分观测值标记为“不可靠”并附注非标准定义说明;不同国家与年份间数据来源的变更(如调查方法修订)造成时间序列断点,需通过注释字段追溯;此外,多个来源并存时需依赖ILO人工筛选“最佳来源”,可能引入主观偏差,且高精度地理维度(如残疾人群体的城乡细分)样本量有限,影响统计推断的稳健性。
常用场景
经典使用场景
该数据集汇集了欧洲30个国家2002至2025年间按性别、城乡区域和残疾状况划分的失业人数(单位:千人)的年度观测数据,是劳动经济学、社会政策与人口统计学领域中进行跨国比较和时间序列分析的理想素材。研究者可借助其多维分类变量,探究不同性别、居住地类型及残疾身份在失业率上的结构性差异,或通过时序建模揭示欧洲劳动力市场波动的长期趋势与周期性特征。
解决学术问题
数据集的精细分层设计有效回应了长期困扰劳动经济学研究的若干核心问题:如何量化并分解性别、城乡和残疾状态这三重因素对失业率的叠加影响,以及这些影响在欧洲各国间的异质性表现。通过系统整合ILOSTAT按照国际劳工统计学家会议定义进行统一化的微观调查数据,该数据集为探讨劳动力市场中的弱势群体就业歧视、区域发展不平衡以及社会包容政策的有效性提供了坚实的数据基础,显著推动了关于公平就业与福祉的量化评估。
实际应用
在实际应用中,该数据集可服务于欧盟及各成员国的劳动就业部门与社会保障机构,用于构建失业风险预警模型和就业促进政策的绩效评估。基于多维分类的失业数据还能辅助企业人力资源部门进行区域化的人才供给预测与多元化招聘策略制定。此外,非政府组织与社会研究智库可据此分析政策干预对残障人士及农村女性等特定群体的实际效果,为倡导更具包容性的劳动立法提供经验证据。
数据集最近研究
最新研究方向
在劳动经济学与交叉不平等研究领域,该数据集为探究性别、城乡地域及残疾状况三重维度下的失业差异提供了稀缺的纵向微观数据支撑。前沿方向聚焦于利用时间序列预测模型与面板回归方法,量化后疫情时代欧洲劳动力市场中弱势群体的结构性失业风险。结合欧盟推动的包容性就业政策与2030年可持续发展议程,研究者可借此揭示残疾人口与农村女性等边缘化人群在劳动力市场中的系统性排斥机制。其长周期、多国别的粒度数据为评估“公平转型”与福祉干预成效铺就了经验基础,已成为检验制度包容性理论的重要实证资产。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务