electricsheepafrica/africa-ilo-ear-thra-sex-nb-average-hourly-earnings-of-tourism-sector-employee
收藏资源简介:
该数据集名为“按性别划分的旅游部门员工平均小时收入(本地货币)| 非洲(ILOSTAT)”,由Electric Sheep Africa重新打包发布。它包含165个观测值,覆盖17个非洲国家(如埃及、赞比亚、安哥拉等),时间跨度为2012年至2024年。数据聚焦于一个核心指标:EAR_THRA_SEX_NB,即按性别划分的旅游部门员工平均小时收入(以本地货币计)。数据来源于国际劳工组织(ILO)的ILOSTAT数据库,通过REST API获取,并经过过滤仅包含非洲国家。数据集采用表格格式,包含列如ref_area(国家代码)、indicator(指标代码)、sex(性别分类)、time(年份)、obs_value(观测值)等,支持表格分类、回归和时间序列预测任务。数据质量方面,为年度频率,ILO选择“最佳来源”处理多源数据,且分解列(如性别)仅在指标发布时非空。用途包括使用Hugging Face的datasets库加载、按国家或时间序列过滤、以及数据透视分析。数据集遵循cc-by-4.0许可,使用时需引用ILO原始来源和Electric Sheep Africa的重新打包。
This dataset is named *Average Hourly Earnings of Tourism Sector Employees by Gender (Local Currency) | Africa (ILOSTAT)*, and is republished by Electric Sheep Africa. It contains 165 observations covering 17 African countries including Egypt, Zambia, Angola and others, with a time span from 2012 to 2024. The dataset focuses on a core indicator: EAR_THRA_SEX_NB, which refers to the average hourly earnings of tourism sector employees by gender, denominated in local currency. The data is sourced from the International Labour Organization (ILO) ILOSTAT database, obtained via REST API, and filtered to only include African countries. In tabular format, the dataset includes columns such as ref_area (country code), indicator (indicator code), sex (gender category), time (year), obs_value (observed value) and others, supporting tasks including tabular classification, regression and time series forecasting. Regarding data quality, the data has an annual frequency; the ILO uses the "best source" approach to handle multi-source data, and disaggregated columns such as gender are only non-empty when the corresponding indicator is published. Possible applications include loading the dataset via the Hugging Face datasets library, filtering by country or time series, and conducting pivot table analysis. This dataset is licensed under CC BY 4.0, and users must cite both the original ILO source and the republishing work of Electric Sheep Africa when using the dataset.




