遇见数据集

electricsheepeurope/europe-ilo-une-tune-sex-eco-nb-unemployment-of-previously-employed-persons-by-sex

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - unemployment - ilo - labour - employment pretty_name: "Unemployment of previously employed persons by sex and former economic activity (thousands | Europe (ILOSTAT)" --- # Unemployment of previously employed persons by sex and former economic activity (thousands | Europe (ILOSTAT) 🇪🇺 **96,748 observations** · **39 Europe countries** · **1970–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-96,748-blue) ![countries](https://img.shields.io/badge/countries-39-green) ![years](https://img.shields.io/badge/years-1970–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **96,748 observations** of `Unemployment` data across **39 Europe countries**, spanning **1970–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_ECO_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Unemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=UNE_TUNE_SEX_ECO_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 39 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `ESP` | 3,656 | 1971 | 2025 | | `BEL` | 3,647 | 1970 | 2024 | | `FIN` | 3,463 | 1976 | 2024 | | `ITA` | 3,317 | 1992 | 2024 | | `GBR` | 3,293 | 1993 | 2025 | | `AUT` | 3,229 | 1983 | 2025 | | `GRC` | 3,222 | 1981 | 2024 | | `FRA` | 3,188 | 1993 | 2024 | | `PRT` | 3,172 | 1974 | 2024 | | `DEU` | 3,153 | 1993 | 2024 | | `HUN` | 3,097 | 1992 | 2024 | | `CZE` | 3,066 | 1993 | 2024 | | `CHE` | 3,038 | 1991 | 2025 | | `DNK` | 2,948 | 1983 | 2024 | | `ROU` | 2,938 | 1994 | 2024 | | ... | _24 more countries_ | | | ## Indicators (sample) - `UNE_TUNE_SEX_ECO_NB` — Unemployment of previously employed persons by sex and former economic activity (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:480` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `UNE_TUNE_SEX_ECO_NB` | | `indicator.label` | `string` | Indicator name in English | `Unemployment of previously employed p…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `ECO_SECTOR_TOTAL` | | `classif1.label` | `string` | — | `Economic activity (Broad sector): Total` | | `time` | `int64` | Observation year | `2024` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `24.601` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C5:3051` | | `note_classif.label` | `string` | — | `Nonstandard economic activity: Includ…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-une-tune-sex-eco-nb-unemployment-of-previously-employed-persons-by-sex") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "UNE_TUNE_SEX_ECO_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="UNE_TUNE_SEX_ECO_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "UNE_TUNE_SEX_ECO_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_une_tune_sex_eco_nb_unemployment_of_previously_employed_persons_by_sex_2025, title = {Unemployment of previously employed persons by sex and former economic activity (thousands | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_ECO_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-une-tune-sex-eco-nb-unemployment-of-previously-employed-persons-by-sex}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-27 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=UNE_TUNE_SEX_ECO_NB_

This dataset contains unemployment data from the International Labour Organization (ILO) ILOSTAT database, specifically the indicator Unemployment of previously employed persons by sex and former economic activity (thousands) for Europe. It includes 96,748 observations across 39 European countries, spanning the years 1970 to 2025. The data is in tabular format and suitable for tabular classification, tabular regression, and time-series forecasting tasks. Data is sourced directly from the ILOSTAT REST API and harmonized by the ILOs Department of Statistics using International Conference of Labour Statisticians (ICLS) definitions. The dataset features columns such as country code, source, indicator, sex disaggregation, year, observed value, and includes notes on data quality and usage examples.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-une-tune-sex-eco-nb-unemployment-of-previously-employed-persons-by-sex 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的ILOSTAT中央统计数据库,聚焦于欧洲地区先前就业者的失业状况,并按性别与先前经济活动进行细致划分。数据通过ILOSTAT REST API直接拉取,并筛选出欧洲39个国家的ISO3代码区域。ILOSTAT依据国际劳工统计学家会议(ICLS)的定义对原始调查微观数据进行协调与标准化,确保数据的一致性与可比性。所有观测值均附有来源标签(source.label),以便追溯数据出处。数据集共包含96,748条观察记录,时间跨度从1970年至2025年,涵盖单一核心指标,并以Parquet格式打包,便于机器学习与数据分析场景直接调用。
特点
该数据集最显著的特点在于其多维度的细粒度解构与权威性。数据按照性别(总计、男性、女性)和先前经济活动(如总经济部门)进行交叉分类,为研究失业结构性特征提供了宝贵素材。时间跨度长达半个世纪,覆盖39个欧洲国家,形成了丰富的时间序列与面板数据结构。数据来源多元,包括劳动力调查、行政记录等,且ILO对多源数据进行了最佳来源选择与质量标注,观察值附有可靠性状态标志(如provisional、unreliable),便于用户进行数据质量评估。这些特性使其成为分析欧洲劳动力市场动态、性别差异以及经济结构变迁的理想资源。
使用方法
该数据集已被封装为HuggingFace Datasets格式,使用者可通过`load_dataset`函数一行代码加载至Python环境,并便捷地转换为Pandas DataFrame进行后续操作。用户可依据`ref_area`字段对特定国家进行筛选,例如提取德国(DEU)的数据进行深入分析。针对时间序列分析,可对`indicator`与`time`字段进行排序与可视化,揭示失业趋势。此外,通过透视表(pivot_table)操作,能轻松构建以时间为行、国家为列的观测值矩阵,适用于面板数据回归或跨国家比较研究。数据支持表格分类、回归及时间序列预测等多种任务,为计量经济学与机器学习研究提供了灵活且高质量的基础数据支撑。
背景与挑战
背景概述
该数据集源自国际劳工组织(ILO)的ILOSTAT统计数据库,于2025年由Electric Sheep Europe团队重新打包发布,聚焦欧洲39个国家1970至2025年间先前就业人员的失业情况,包含96,748条观测记录。作为全球劳动统计的权威来源,ILO通过协调各国劳动力调查、家庭收入调查及行政记录,系统汇编了涵盖就业、失业、工资等领域的指标。该数据集的核心研究问题在于揭示欧洲区域不同性别及先前经济活动中失业人口的结构性特征,为劳动经济学、社会政策及跨国比较研究提供可靠的数据基础。其影响力体现在为时间序列分析与表格回归任务提供标准化、机器可读的开放式数据资源,推动了欧洲劳动力市场动态的量化研究。
当前挑战
数据集面临的挑战源于多层面。在领域问题层面,失业统计需精准捕捉经济转型(如2008年金融危机及新冠疫情期间)导致的劳动市场波动,而不同国家收集指标时的定义差异(如失业认定标准)与统计方法变更(如“方法修订”标注)可能干扰时间序列的一致性。在构建过程中,数据源于39个国家的多源调查(如劳动力调查),需要解决来源间数据整合难题,包括处理缺失值、识别“不可靠”观测状态及统一因分类维度(性别、经济活动)导致的细粒度分解。此外,年度频率限制了对短期波动的分析能力,而原始数据中部分月度与季度序列未被纳入,进一步制约了高频建模的精度。
常用场景
经典使用场景
在劳动经济学与泛欧社会政策研究领域,欧洲劳动力市场的结构性变迁始终是学界关注的焦点。europe-ilo-une-tune-sex-eco-nb数据集凭借其涵盖39个欧洲国家、横跨1970年至2025年长达半个多世纪的丰富观测数据,成为探究先前就业人员失业率时空演变的经典资源。研究者常利用该数据集开展时间序列分析与面板数据建模,通过按性别与前经济部门划分的失业人数指标,精细描绘不同人口群体在特定产业周期中的就业脆弱性。数据集提供的ISO国家代码、ILOSTAT权威指标以及详尽的分类维度,使得跨国比较与历史趋势挖掘得以高效实现,为构建解释欧洲失业动态的计量经济学模型奠定了坚实的数据基础。
实际应用
在现实政策与商业决策层面,该数据集具备广泛的应用潜力。各国劳动与社会保障部门可借助历史失业结构化数据,构建早期预警模型以识别特定行业或地区的就业风险聚集,进而优化公共就业服务的资源配置。国际组织如欧盟委员会与国际货币基金组织在评估成员国劳动力市场韧性、监测欧洲就业战略实施进展时,可直接将此数据集纳入其政策模拟与情景分析框架。此外,人力资源咨询公司与金融机构亦可利用该数据洞察不同性别与产业背景下的劳动力供给变化,辅助判断区域投资环境与潜在的人力资本瓶颈,从而更精准地制定跨境招募策略或风险评估模型。
衍生相关工作
围绕这一高价值数据集,学术界与工业界催生了一系列衍生性与拓展性工作。在数据工程层面,Electric Sheep Europe团队建立的自动化管道实现了ILOSTAT原始数据的定期抓取、模式规范化与Parquet格式打包,推动了欧洲官方统计数据的机器学习就绪水平。在分析范式层面,该数据集常与劳动力参与率、工资分布、临时就业比例等互补指标进行融合,构建多维度的欧洲劳动力市场健康状况评估体系。此外,基于该数据的时间序列特征,已有研究将其纳入贝叶斯结构时间序列模型与深度学习预测框架中,用于对欧元区失业率的短期预测与反事实政策评估,这些工作进一步扩展了微观劳动统计在宏观计量与因果推断中的应用边界。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务