electricsheepafrica/africa-ilo-emp-temp-sex-edu-geo-nb-employment-by-sex-education-and-rural-urban-areas
收藏资源简介:
--- license: cc-by-4.0 language: - en task_categories: - tabular-classification - tabular-regression - time-series-forecasting multilinguality: monolingual size_categories: - 10K<n<100K tags: - tabular - africa - ilostat - employment - ilo - labour pretty_name: "Employment by sex, education and rural / urban areas (thousands) | Africa (ILOSTAT)" --- # Employment by sex, education and rural / urban areas (thousands) | Africa (ILOSTAT) 🌍 **29,281 observations** · **45 Africa countries** · **1994–2025** · *Repackaged by [Electric Sheep Africa](https://huggingface.co/electricsheepafrica)*      ## TL;DR This dataset contains **29,281 observations** of `Employment` data across **45 Africa countries**, spanning **1994–2025**, covering **1 distinct indicators**. ## About the source **ILOSTAT** is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment, unemployment, wages, working time, child labour, informal economy, social protection, occupational injuries, and SDG decent work targets — drawing on national labour force surveys, household income surveys, establishment surveys, and administrative records. Coverage spans 200+ economies, with the ILO's Department of Statistics responsible for harmonisation. - **Source:** [ILOSTAT](https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_TEMP_SEX_EDU_GEO_NB) - **Publisher:** International Labour Organization (ILO) - **License:** [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/) - **Topic:** Employment ## Methodology Data pulled directly from the ILOSTAT REST API at `https://rplumber.ilo.org/data/indicator?id=EMP_TEMP_SEX_EDU_GEO_NB` and filtered to Africa ISO3 country codes. ILOSTAT harmonises raw survey microdata using ICLS (International Conference of Labour Statisticians) definitions; sources are flagged in the `source.label` column for traceability. ## Geographic coverage 45 Africa countries · top rows shown below, sorted by row count: | Country | Rows | First year | Last year | |---------|-----:|-----------:|----------:| | `ZAF` | 2,495 | 2008 | 2024 | | `EGY` | 2,085 | 2008 | 2024 | | `AGO` | 1,462 | 2004 | 2025 | | `GHA` | 1,451 | 2000 | 2024 | | `MLI` | 1,362 | 2013 | 2024 | | `ZMB` | 1,228 | 2015 | 2024 | | `RWA` | 1,224 | 2014 | 2025 | | `TUN` | 1,120 | 2006 | 2023 | | `SEN` | 1,036 | 2011 | 2024 | | `UGA` | 952 | 2010 | 2021 | | `ZWE` | 922 | 2011 | 2024 | | `TZA` | 882 | 2001 | 2020 | | `TGO` | 848 | 2006 | 2022 | | `NAM` | 784 | 1994 | 2018 | | `KEN` | 780 | 1999 | 2022 | | ... | _30 more countries_ | | | ## Indicators (sample) - `EMP_TEMP_SEX_EDU_GEO_NB` — Employment by sex, education and rural / urban areas (thousands) ## Schema | Column | Type | Description | Example | |--------|------|-------------|---------| | `ref_area` | `string` | ISO 3166-1 alpha-3 country code | `AGO` | | `ref_area.label` | `string` | Country name in English | `Angola` | | `source` | `string` | ILOSTAT source code (e.g. labour force survey) | `BA:13951` | | `source.label` | `string` | Source name in English | `LFS - Employment Survey` | | `indicator` | `string` | ILOSTAT indicator code | `EMP_TEMP_SEX_EDU_GEO_NB` | | `indicator.label` | `string` | Indicator name in English | `Employment by sex, education and rura…` | | `sex` | `string` | Disaggregation by sex (SEX_T = total, SEX_M = male, SEX_F = female) | `SEX_T` | | `sex.label` | `string` | — | `Total` | | `classif1` | `string` | First classification variable (age, education, status, etc.) | `EDU_AGGREGATE_TOTAL` | | `classif1.label` | `string` | — | `Education (Aggregate levels): Total` | | `classif2` | `string` | Second classification variable where applicable | `GEO_COV_NAT` | | `classif2.label` | `string` | — | `Area type: National` | | `time` | `int64` | Observation year | `2025` | | `obs_value` | `float64` | Observed indicator value (unit varies — see indicator definition) | `13984.984` | | `obs_status` | `string` | Observation status flag (e.g. provisional, unreliable) | `U` | | `obs_status.label` | `string` | — | `Unreliable` | | `note_classif` | `string` | — | `C3:2620` | | `note_classif.label` | `string` | — | `Nonstandard education level: Includin…` | | `note_indicator` | `string` | — | `I11:264` | | `note_indicator.label` | `string` | — | `Break in series: Methodology revised` | | `note_source` | `string` | — | `R1:3513` | | `note_source.label` | `string` | — | `Repository: ILO-STATISTICS - Micro da…` | ## Disaggregation dimensions The following columns provide disaggregation dimensions: - **`sex`** (3 unique values): `SEX_T`, `SEX_M`, `SEX_F` ## Data quality & caveats - Data is annual frequency. Some indicators also publish monthly or quarterly series — those are not included here. - When an indicator has multiple sources for the same country×year, the ILO-selected 'best source' is used. - Disaggregation columns (`sex`, `classif1`, `classif2`) are non-null only when the indicator publishes that breakdown. ## Usage ```python from datasets import load_dataset ds = load_dataset("electricsheepafrica/africa-ilo-emp-temp-sex-edu-geo-nb-employment-by-sex-education-and-rural-urban-areas") df = ds["train"].to_pandas() print(df.head()) ``` ### Filter to one country ```python kenya = df[df["ref_area"] == "KEN"] ``` ### Time-series for a single indicator ```python sample = (df[df["indicator"] == "EMP_TEMP_SEX_EDU_GEO_NB"] .sort_values("time")) sample.plot(x="time", y="obs_value", title="EMP_TEMP_SEX_EDU_GEO_NB") ``` ### Pivot to country × year matrix ```python matrix = (df[df["indicator"] == "EMP_TEMP_SEX_EDU_GEO_NB"] .pivot_table(index="time", columns="ref_area", values="obs_value")) print(matrix.tail()) ``` ## Citation ```bibtex @misc{africa_ilo_emp_temp_sex_edu_geo_nb_employment_by_sex_education_and_rural_urban_areas_2025, title = {Employment by sex, education and rural / urban areas (thousands) | Africa (ILOSTAT)}, author = {International Labour Organization (ILO)}, year = {2025}, url = {https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_TEMP_SEX_EDU_GEO_NB}, publisher = {HuggingFace Datasets, repackaged by Electric Sheep Africa}, howpublished = {\url{https://huggingface.co/datasets/electricsheepafrica/africa-ilo-emp-temp-sex-edu-geo-nb-employment-by-sex-education-and-rural-urban-areas}} } ``` ## License Released under [cc-by-4.0](https://creativecommons.org/licenses/by/4.0/). Original data © International Labour Organization (ILO). When using this dataset, please cite both the original source above and the Electric Sheep Africa repackaging. ## About Electric Sheep Electric Sheep Africa is part of the Electric Sheep mission: a unified, ML-ready data layer for Africa on HuggingFace. We pull data from authoritative open sources, normalize the schemas, package as Parquet, and publish with consistent dataset cards so researchers and developers can use `load_dataset()` to start working in seconds. Browse the full collection: [huggingface.co/electricsheepafrica](https://huggingface.co/electricsheepafrica) --- _Provenance: ingested 2026-05-26 via the Electric Sheep pipeline. Source URL: https://www.ilo.org/shinyapps/bulkexplorer/?id=EMP_TEMP_SEX_EDU_GEO_NB_
This dataset contains 29,281 observations of Employment data across 45 Africa countries, spanning 1994–2025, covering 1 distinct indicator. The data is sourced from ILOSTAT, the International Labour Organizations central statistics database, obtained via its REST API and filtered to Africa ISO3 country codes. The key indicator is EMP_TEMP_SEX_EDU_GEO_NB, which represents employment figures disaggregated by sex, education level, and rural/urban areas (in thousands). The dataset includes columns such as country code, country name, data source, indicator code, indicator label, sex classification (total, male, female), education classification, area type, observation year, observed value, observation status, and related notes. Data is annual frequency and harmonized using ICLS (International Conference of Labour Statisticians) definitions. For data quality, when multiple sources exist for the same country and year, the ILO-selected best source is used. This dataset is suitable for tabular classification, regression, and time-series forecasting tasks, providing ML-ready employment data for Africa.




