遇见数据集

electricsheepeurope/europe-ilo-ees-tees-sex-edu-geo-nb-employees-by-sex-education-and-rural-urban-areas-t

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - europe - ilostat - employees - ilo - labour - employment pretty_name: "Employees by sex, education and rural / urban areas (thousands) | Europe (ILOSTAT)" --- # Employees by sex, education and rural / urban areas (thousands) | Europe (ILOSTAT) 🇪🇺 **83,154 observations** · **37 Europe countries** · **1987–2025** · *Repackaged by [Electric Sheep Europe](https://huggingface.co/electricsheepeurope)* ![rows](https://img.shields.io/badge/rows-83,154-blue) ![countries](https://img.shields.io/badge/countries-37-green) ![years](https://img.shields.io/badge/years-1987–2025-orange) ![indicators](https://img.shields.io/badge/indicators-1-purple) ![license](https://img.shields.io/badge/license-cc-by-4.0-lightgrey) ## TL;DR This dataset contains **83,154 observations** of `Employees` data across **37 Europe countries**, spanning **1987–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_EDU_GEO_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Employees ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EES_TEES_SEX_EDU_GEO_NB` and filtered to Europe ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 37 Europe countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `GRC` | 5,051 | 1987 | 2025 | | `FRA` | 3,276 | 1993 | 2024 | | `IRL` | 3,142 | 1993 | 2024 | | `NLD` | 3,017 | 1996 | 2024 | | `ITA` | 3,000 | 1992 | 2024 | | `BEL` | 2,967 | 1992 | 2024 | | `PRT` | 2,964 | 1992 | 2024 | | `ESP` | 2,911 | 1992 | 2024 | | `DNK` | 2,896 | 1992 | 2024 | | `DEU` | 2,891 | 1992 | 2024 | | `SWE` | 2,837 | 1995 | 2024 | | `LUX` | 2,828 | 1997 | 2024 | | `AUT` | 2,700 | 1995 | 2025 | | `FIN` | 2,642 | 1995 | 2024 | | `SRB` | 2,546 | 2007 | 2025 | | ... | _22 more countries_ | | | ## Indicators (sample) - `EES_TEES_SEX_EDU_GEO_NB` — Employees by sex, education and rural / urban areas (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `ALB` | | `ref_area.label` | `string` | Country name in English | `Albania` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BB:7401` | | `source.label` | `string` | Source name in English | `HIES - Living Standards Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EES_TEES_SEX_EDU_GEO_NB` | | `indicator.label` | `string` | Indicator name in English | `Employees by sex, education and rural…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `classif2` | `string` | Second classification variable where applicable | `GEO_COV_NAT` | | `classif2.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2012` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `421.776` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:5578` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepeurope/europe-ilo-ees-tees-sex-edu-geo-nb-employees-by-sex-education-and-rural-urban-areas-t") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python germany = df[df["ref_area"] == "DEU"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EES_TEES_SEX_EDU_GEO_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EES_TEES_SEX_EDU_GEO_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EES_TEES_SEX_EDU_GEO_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{europe_ilo_ees_tees_sex_edu_geo_nb_employees_by_sex_education_and_rural_urban_areas_t_2025, title = {Employees by sex, education and rural / urban areas (thousands) | Europe (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_EDU_GEO_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Europe}, howpublished = {\url{https://huggingface.co/datasets/electricsheepeurope/europe-ilo-ees-tees-sex-edu-geo-nb-employees-by-sex-education-and-rural-urban-areas-t}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Europe repackaging. ## About Electric Sheep Electric Sheep Europe is part of the Electric Sheep mission: a unified, ML-ready data layer for Europe on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepeurope](https://huggingface.co/electricsheepeurope) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EES_TEES_SEX_EDU_GEO_NB_

This dataset contains 83,154 observations of employees data across 37 European countries, spanning from 1987 to 2025, with the core indicator EES_TEES_SEX_EDU_GEO_NB, which represents employees disaggregated by sex, education level, and rural/urban areas (in thousands). The data is sourced from the International Labour Organization (ILO) ILOSTAT database, retrieved via API and filtered for European ISO3 country codes. It includes multiple dimension columns such as country code, country name, data source, indicator code, sex classification, education classification, rural/urban classification, observation year, observed value, and status flags, along with data quality caveats and usage examples, suitable for tabular classification, regression, and time-series forecasting tasks.

提供机构:
electricsheepeurope
搜集汇总
数据集介绍
electricsheepeurope/europe-ilo-ees-tees-sex-edu-geo-nb-employees-by-sex-education-and-rural-urban-areas-t 数据集图片
构建方式
该数据集依托国际劳工组织(ILO)的ILOSTAT中央统计数据库构建,通过其REST API接口直接提取指标EES_TEES_SEX_EDU_GEO_NB的原始记录,并依据ISO 3166-1 alpha-3国家代码筛选欧洲区域数据。源数据由ILO统计部门依据国际劳工统计学家会议(ICLS)定义对各国劳动力调查、家庭收入调查及行政记录等微观数据进行标准化整合,最终由Electric Sheep Europe重新封装为列式存储格式,涵盖1987至2025年间37个欧洲国家的83,154条观测记录,并保留来源标签以确保可追溯性。
使用方法
研究者可通过HuggingFace的datasets库以单行代码加载数据集,并直接转换为Pandas数据框进行灵活操作。例如,可筛选特定国家(如德国DEU)进行国别分析,或提取单一指标构建时间序列图以观察趋势演变。此外,用户可利用透视表功能将数据重塑为国家与年份的交叉矩阵,便于开展面板数据分析或跨国比较研究。数据集遵循CC-BY-4.0许可协议,使用时需同时引用ILO原始来源与Electric Sheep Europe的再封装工作。
背景与挑战
背景概述
国际劳工组织(ILO)长期致力于全球劳动力市场统计数据的标准化与传播,其维护的ILOSTAT数据库是劳动经济学领域最具权威性的跨国数据来源之一。在此背景下,Electric Sheep Europe对ILOSTAT原始数据进行了系统性重封装,于2025年发布了该数据集,覆盖1987至2025年间37个欧洲国家约83,154条雇员观测记录,按性别、教育程度及城乡区域进行多维 disaggregation。该数据集的核心研究问题在于揭示欧洲劳动力市场中雇员分布的性别差异、教育回报与城乡分化格局,为劳动经济学、教育政策评估及区域发展研究提供了坚实的数据基础。其影响力在于降低了跨国比较研究的门槛,使研究者能够借助统一schema快速开展面板分析与时间序列预测。
当前挑战
该数据集所应对的领域问题在于,如何在跨国、跨时间的异质性语境下实现雇员统计指标的可比性与一致性,这要求对各国差异化的劳动力调查口径进行复杂的事后协调。在构建过程中,主要挑战体现在三方面:其一,数据来源的多样性导致同一国家同年份可能存在多个调查来源,ILO虽采用‘最佳来源’原则进行筛选,但来源间的口径差异仍可能引入测量误差;其二,部分观测值存在可靠性标记(如provisional或unreliable)以及序列断点(如方法论修订),这对时间序列建模的连续性构成潜在威胁;其三,教育与城乡分类体系在不同国家间的非标准化问题,使得深层次的分组比较面临语义对齐的困难,需要研究者在分析时审慎处理分类变量的映射与聚合。
常用场景
经典使用场景
在劳动经济学与性别研究的量化分析中,该数据集构成了一项基础性观测资源。研究者依托其涵盖37个欧洲国家、1987年至2025年间逾八万条记录,按性别、教育程度与城乡地域对雇员人数进行细粒度分解,得以开展面板数据回归、时间序列预测与跨国比较分析。尤为典型的是,以性别与教育维度交叉考察就业结构变迁,可揭示不同区域劳动力市场中的结构性差异与趋同趋势。该数据集亦常用于训练表格分类与回归模型,以预测特定人群的就业水平。
解决学术问题
该数据集有效回应了劳动经济学中长期存在的若干研究难题,包括性别就业差距的教育梯度演化、城乡劳动力市场分割的跨国模式识别,以及教育扩张对女性就业参与率的异质性影响。借由ILO统一协调的ICLS定义,它缓解了以往跨国比较中统计口径不一致的痼疾,使研究者能够在标准化框架下检验制度变迁与政策干预的效应。其长时序跨度亦为识别结构性断点与周期性波动提供了实证基础,对理解欧洲一体化的劳动力市场后果具有重要学术意义。
实际应用
在政策制定与实务领域,该数据集为国际组织、国家统计部门及智库提供了监测欧洲就业结构变化的量化依据。决策者可据此评估教育政策与区域发展策略在促进性别平等就业方面的成效,识别城乡就业机会失衡的突出区域,并制定有针对性的劳动力市场干预措施。企业与国际机构亦可利用该数据研判不同教育层次劳动力的供给格局,优化人力资源配置与跨国投资决策。此外,该数据集支持构建可视化仪表板,辅助公众理解就业趋势。
数据集最近研究
最新研究方向
在全球劳动力市场结构性转型与教育公平议题交织的背景下,该数据集以性别、教育程度及城乡地域的三重交叉分层为切入点,为解析欧洲各国就业结构的异质性提供了精细化的微观证据。近期研究前沿聚焦于运用面板数据与机器学习方法,探究教育扩张对性别就业差距的收敛效应,以及城乡分割如何重塑高学历女性的劳动参与轨迹。伴随ILO体面劳动议程与欧盟性别平等战略的政策热点,此类时序跨国数据被广泛用于评估结构性改革对弱势群体就业的差异化影响,其价值在于以可追溯的标准化指标支撑因果推断与预测建模,进而为弥合区域发展鸿沟提供实证依据。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务