electricsheepasia/asia-ilo-ear-ehra-sex-ind-cur-nb-average-hourly-earnings-of-employees-by-ilo-sector
收藏资源简介:
该数据集名为“按ILO部门、性别和货币划分的员工平均小时收入 | 亚洲(ILOSTAT)”,包含29,390个观测值,覆盖22个亚洲国家(如土耳其、越南、柬埔寨等),时间跨度为2005年至2025年。数据来源于国际劳工组织(ILO)的ILOSTAT数据库,这是一个全球劳动力统计的主要来源,通过标准化调查数据(使用国际劳工统计学家会议定义)收集。数据集包含一个核心指标:EAR_EHRA_SEX_IND_CUR_NB,即按ILO部门、性别和货币划分的员工平均小时收入。数据以表格形式组织,包括列如国家代码、国家名称、数据来源、指标代码、性别分类(总计、男性、女性等)、时间年份、观测值(收入数值)以及数据质量标志(如不可靠)。数据为年度频率,经过ILO的“最佳来源”选择处理,并提供了分类维度(如性别)以支持详细分析。该数据集由Electric Sheep Asia重新打包,以机器学习友好的格式(Parquet)发布,便于使用Hugging Face的datasets库直接加载和分析。
The dataset, titled Average hourly earnings of employees by ILO sector, sex and currency | Asia (ILOSTAT), comprises 29,390 observations across 22 Asian countries (e.g., Turkey, Vietnam, Cambodia), spanning the years 2005 to 2025. It is sourced from the ILOSTAT database of the International Labour Organization (ILO), a leading global repository for labour statistics, which harmonizes raw survey microdata using International Conference of Labour Statisticians (ICLS) definitions. The dataset focuses on a single indicator: EAR_EHRA_SEX_IND_CUR_NB, representing average hourly earnings of employees disaggregated by ILO sector, sex, and currency. Data is structured in tabular format with columns including country code, country name, source code, indicator code, sex classification (total, male, female, etc.), time year, observed value (earnings), and data quality flags (e.g., unreliable). It is annual frequency data, with ILOs best source selection applied for duplicate entries, and includes disaggregation dimensions such as sex for granular analysis. Repackaged by Electric Sheep Asia in a machine-learning-ready format (Parquet), it allows easy loading and analysis via Hugging Faces datasets library.




