electricsheepeurope/europe-ilo-ees-tees-sex-mts-nb-employees-by-sex-and-marital-status-thousands
收藏资源简介:
该数据集名为按性别和婚姻状况划分的雇员(千人)| 欧洲(ILOSTAT),是一个表格型数据集,专注于时间序列预测、表格分类和回归任务。它包含14,402条观测数据,覆盖39个欧洲国家,时间跨度为1983年至2025年。数据集的核心指标是EES_TEES_SEX_MTS_NB,即按性别和婚姻状况划分的雇员数量(以千人为单位)。数据来源于国际劳工组织(ILO)的ILOSTAT数据库,该数据库是全球劳动力统计的主要来源,通过国家劳动力调查、家庭收入调查、机构调查和行政记录收集数据。数据集经过ILO使用国际劳工统计学家会议(ICLS)定义进行标准化处理,并通过Electric Sheep Europe重新打包,以提供一致的ML就绪格式。数据模式包括国家代码、国家名称、数据来源、指标代码、指标标签、性别分类(总计、男性、女性)、婚姻状况分类、观测年份、观测值、观测状态标志以及相关注释。数据按性别维度进行细分,提供SEX_T(总计)、SEX_M(男性)和SEX_F(女性)三个唯一值。数据质量方面,数据为年度频率,ILO会选择最佳来源处理同一国家×年份的多个数据源,且细分列仅在指标发布该细分时非空。数据集旨在用于劳动力市场分析、经济研究和机器学习建模,支持按国家筛选、时间序列分析和数据透视等操作。
The dataset is named Employees by sex and marital status (thousands) | Europe (ILOSTAT) and is a tabular dataset focused on time-series forecasting, tabular classification, and regression tasks. It contains 14,402 observations across 39 European countries, spanning from 1983 to 2025. The core indicator is EES_TEES_SEX_MTS_NB, which represents Employees by sex and marital status (in thousands). The data is sourced from the International Labour Organization (ILO)s ILOSTAT database, a leading global source for labour statistics, compiled from national labour force surveys, household income surveys, establishment surveys, and administrative records. The dataset is harmonized by ILO using International Conference of Labour Statisticians (ICLS) definitions and repackaged by Electric Sheep Europe to provide a consistent ML-ready format. The schema includes columns for country code, country name, data source, indicator code, indicator label, sex disaggregation (total, male, female), marital status classification, observation year, observed value, observation status flags, and relevant notes. The data is disaggregated by sex, with three unique values: SEX_T (total), SEX_M (male), and SEX_F (female). In terms of data quality, the data is annual in frequency; ILO selects the best source when multiple sources exist for the same country×year, and disaggregation columns are non-null only when the indicator publishes that breakdown. The dataset is intended for labour market analysis, economic research, and machine learning modeling, supporting operations such as filtering by country, time-series analysis, and data pivoting.




