遇见数据集

electricsheepeurope/europe-ilo-emp-xtru-sex-age-nb-time-related-underemployment-by-sex-and-age-thousa

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - time-related-underemployment - ilo - labour - employment pretty_name: "Time-related underemployment by sex and age (thousands) | Europe (ILOSTAT)" --- # Time-related underemployment by sex and age (thousands) | Europe (ILOSTAT) 🇪🇺 **36,685 observations** · **39 Europe countries** · **1991–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-36,685-blue) ![countries](https://img.shields.io/badge/countries-39-green) ![years](https://img.shields.io/badge/years-1991–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **36,685 observations** of `Time-related underemployment` data across **39 Europe countries**, spanning **1991–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_AGE_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Time-related underemployment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EMP_XTRU_SEX_AGE_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 39 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `PRT` | 1,283 | 1998 | 2025 | | `CHE` | 1,265 | 1991 | 2025 | | `GBR` | 1,262 | 1999 | 2025 | | `AUT` | 1,256 | 1998 | 2025 | | `FRA` | 1,254 | 1998 | 2024 | | `ESP` | 1,228 | 1999 | 2025 | | `FIN` | 1,178 | 1999 | 2024 | | `POL` | 1,168 | 2001 | 2025 | | `SWE` | 1,157 | 2000 | 2024 | | `NLD` | 1,145 | 2000 | 2024 | | `ROU` | 1,138 | 1999 | 2024 | | `LTU` | 1,098 | 2001 | 2024 | | `ITA` | 1,098 | 2002 | 2024 | | `BEL` | 1,081 | 1999 | 2024 | | `EST` | 1,081 | 1998 | 2024 | | ... | _24 more countries_ | | | ## Indicators (sample) - `EMP_XTRU_SEX_AGE_NB` — Time-related underemployment by sex and age (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:480` | | `source.label` | `string` | Source name in English | `LFS - Labour Force Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EMP_XTRU_SEX_AGE_NB` | | `indicator.label` | `string` | Indicator name in English | `Time-related underemployment by sex a…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `AGE_YTHADULT_YGE15` | | `classif1.label` | `string` | — | `Age (Youth, adults): 15+` | | `time` | `int64` | Observation year | `2024` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `38.876` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C6:1634` | | `note_classif.label` | `string` | — | `Nonstandard age group: Including ages…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-emp-xtru-sex-age-nb-time-related-underemployment-by-sex-and-age-thousa") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EMP_XTRU_SEX_AGE_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EMP_XTRU_SEX_AGE_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EMP_XTRU_SEX_AGE_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_emp_xtru_sex_age_nb_time_related_underemployment_by_sex_and_age_thousa_2025, title = {Time-related underemployment by sex and age (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_AGE_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-emp-xtru-sex-age-nb-time-related-underemployment-by-sex-and-age-thousa}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_XTRU_SEX_AGE_NB_

This dataset contains 36,685 observations of Time-related underemployment by sex and age (thousands) data across 39 Europe countries, spanning from 1991 to 2025, covering 1 distinct indicator. The data is sourced from the ILOSTAT database of the International Labour Organization (ILO), retrieved via its REST API, and filtered to European ISO3 country codes. It includes the indicator EMP_XTRU_SEX_AGE_NB, which measures time-related underemployment in the labor force, disaggregated by sex (total, male, female) and age groups. The dataset is structured in tabular format with columns such as country code, year, observed value, source, and quality flags, suitable for tasks like tabular classification, regression, and time-series forecasting. It has been repackaged by Electric Sheep Europe for machine learning readiness and is released under the CC-BY-4.0 license.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-emp-xtru-sex-age-nb-time-related-underemployment-by-sex-and-age-thousa 数据集图片
构建方式
该数据集源自国际劳工组织(ILO)的中央统计数据库ILOSTAT,系全球劳动统计领域的权威数据源之一。构建过程以ILOSTAT提供的REST API为技术接口,直接获取指标代码为EMP_XTRU_SEX_AGE_NB的原始记录,并依据ISO 3166-1 alpha-3国家编码筛选出39个欧洲国家。ILOSTAT对来源于各国劳动力调查、家庭收入调查及行政记录等原始微观数据,按照国际劳工统计学家会议(ICLS)定义进行标准化调和,数据来源信息以source.label字段标注,确保了数据的可追溯性与跨国家可比性。
特点
数据集涵盖1991年至2025年共36,685条观测记录,聚焦于欧洲39国的与时间相关的就业不足人口数量(以千人为单位),按性别与年龄维度进行分解。时间序列跨越逾三十年,兼具截面与纵向分析潜力。观测值附有来源标签、观测状态标志及分类注释等元数据字段,可辅助识别数据修订、非标准年龄分组及序列断裂等情形。性别维度包含总计、男性与女性三类取值,年龄分类则提供青年与成年等分组,结构清晰,便于多维透视。
使用方法
研究者可通过HuggingFace datasets库以load_dataset函数直接加载该数据集,并利用to_pandas方法转换为数据框进行后续分析。典型用法包括:按ref_area字段筛选特定国家,如提取德国(DEU)的全部记录;针对单一指标按时间排序以绘制时间序列趋势图;或通过pivot_table方法将数据重塑为国家与年份的交叉矩阵,以考察就业不足人口数量的时空演变格局。该数据集亦适用于表格分类、回归及时间序列预测等机器学习任务的基准测试。
背景与挑战
背景概述
在全球化劳动力市场深度演变与就业质量议题持续升温的背景下,国际劳工组织(ILO)长期致力于构建跨国可比劳动统计体系,ILOSTAT即其核心数据库。Electric Sheep Europe于2025年对ILOSTAT中时间相关不充分就业指标进行标准化重封装,形成覆盖39个欧洲国家、1991至2025年、逾三万六千条观测的专题数据集。该数据集以性别与年龄为关键分层维度,旨在为就业不足的跨国时序比较提供细粒度、可复现的微观面板支撑,对劳动经济学、社会政策评估及可持续发展目标监测具有基础性数据价值。
当前挑战
时间相关不充分就业的测度本身便深嵌于国际劳工统计学家会议定义与各国劳动力调查实践之间的张力之中,涉及工时门槛、收入意愿与搜寻行为的多重判定。数据集构建面临的核心难题包括:各国调查方法更迭导致的序列断裂、非标准年龄分组的广泛存在、部分观测值因样本量不足而被标记为不可靠,以及按性别与年龄交叉分层后细分单元稀疏性加剧。此外,年度频率数据难以刻画季节性波动与短期政策冲击,跨国口径调和过程中信息损失亦不可避免,这些因素共同构成基于该数据集开展严谨推断的实质约束。
常用场景
经典使用场景
在劳动经济学与就业政策研究领域,时间相关就业不足作为衡量劳动力未充分利用的关键指标,长期依赖于跨国可比数据的支撑。该数据集汇聚了欧洲39个国家自1991年至2025年间按性别与年龄分组的就业不足观测值,总计逾三万六千条记录,为研究者构建面板数据模型、开展时序预测与性别年龄差异分析提供了经典的数据基础。其典型使用场景包括基于国家固定效应的面板回归、性别与年龄组的交叉趋势比较,以及利用时间序列方法识别就业不足的季节性与周期性波动。
解决学术问题
该数据集直面劳动统计研究中长期存在的跨国数据碎片化与口径不一致问题。通过ILOSTAT的统一协调框架,原始调查微数据依据国际劳工统计学家会议定义进行标准化处理,有效缓解了因各国劳动力调查方法差异导致的可比性困境。其学术意义在于为检验就业不足的性别差距假说、青年与成年群体脆弱性差异以及经济周期对劳动力利用不足的异质性影响提供了稳健的实证素材,进而推动了劳动力市场松紧程度测度与体面劳动赤字评估的方法论进展。
衍生相关工作
围绕该数据集及其上游ILOSTAT体系,劳动统计与计量经济学领域衍生出若干经典研究脉络。部分工作聚焦于构建就业不足与失业、非自愿兼职之间的关联指标体系,另有研究将此类数据与宏观经济变量耦合以检验奥肯定律在劳动力利用不足维度的适用性。在方法层面,基于该数据的面板协整与动态因子模型分析亦见诸文献,用以提取欧洲劳动力市场共同波动成分并评估政策协调效应。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务