Vulnerabilidad de Empleos a la Inteligencia Artificial en España: Dataset, Metodología y Dashboard Interactivo (v26B / v13)
收藏资源简介:
AI Exposure of Jobs in Spain — Complete Dataset, Methodology & Interactive Dashboard (V26B/v13) This deposit contains the complete dataset, methodology, and interactive visualisation tool for assessing the theoretical exposure of 502 Spanish occupations to artificial intelligence. The analysis covers 22.46 million workers (EPA Q4 2025, INE) and assigns each occupation a calibrated exposure score on a 0–10 scale, cross-referenced with salary data, EU AI Act risk classification, impact typology, unemployment data, gender breakdowns, and 4-dimensional sub-component scores (D/C/F/R) for all 502 occupations. The interactive dashboard is available at: https://alvarodenicolas.com/interactive/empleos-ia/index.html Mirror: https://empleo-ai.netlify.app/ Dataset The core dataset (spain_502_v13_subcomp_complete.json) contains 502 records corresponding to the complete CNO-11 occupational taxonomy (SEPE expansion). Each record includes 34 fields: Field Type Description cno string 4-digit CNO-11 occupation code (zero-padded) nombre string Official occupation name (Spanish) sector string Assigned economic sector (12 categories) empleo integer Estimated employment (4-layer cascade: EPA Q4 2025 + Census 2021 + SEPE 2024; 499 unique values) empleo_v10 integer Previous employment estimate (v10, preserved for traceability) empleo_delta_v11 integer Employment change v10→v11 salario_medio_eur float Estimated mean gross annual salary (EUR), EES 2023 + educational premia + FR/PT proxies (128 unique values; MAPE 4.96%) vulnerabilidad_ia_score float AI exposure score (0–10), calibrated with 5 Spain-specific factors + sub-component rescoring eu_ai_act string EU AI Act risk: "Alto riesgo" (51 occ), "Riesgo limitado" (4), or "Riesgo mínimo" (447) tipo_impacto string Impact typology: "Sustitución", "Híbrido", or "Aumentación" justificacion string 3–4 sentence justification in Spanish census_2021_employed float Census 2021 employment (3-digit structural weight) epa_2digit_empleo float EPA 2-digit subgroup total (hard constraint) sepe_contracts_2024 integer SEPE 2024 registered contracts (4-digit weight) sepe_contracts_hombres integer SEPE contracts — male sepe_contracts_mujeres integer SEPE contracts — female sepe_parados_dic2024 integer Registered unemployed, December 2024 sepe_parados_hombres_dic2024 integer Registered unemployed Dec 2024 — male sepe_parados_mujeres_dic2024 integer Registered unemployed Dec 2024 — female sepe_parados_jun2024 integer Registered unemployed, June 2024 (seasonal comparison) employment_method string Chain of layers used for employment estimation employment_confidence string "A" (Census+SEPE chain, 382 occ) or "B" (SEPE-synthetic, 120 occ) score_v9 float Previous score (v9, preserved for traceability) rescore_method string "sub-component-4D-rescaled" (177 occ) or "sub-component-4D-informative" (325 occ) rescore_range string Linear rescaling range applied (e.g. "[3.0, 8.5]") rescore_cluster string Original cluster identifier ("7.0", "5.0", or "informative") rescore_D float Sub-component: Digitalisation of tasks (0–10) rescore_C float Sub-component: Current AI capability (0–10) rescore_F float Sub-component: Physical/presence barrier (0–10) rescore_R float Sub-component: Regulatory/institutional friction (0–10) rescore_n_models integer Number of independent models scoring sub-components (2) rescore_raw_avg float Average raw score from sub-component formula rescore_formula string Formula used: "(D×C/10)×(1−F/20)×(1−R/20)" flag_divergencia_gt2 boolean True if |formula score − holistic score| > 2.0 points (158 of 325 informative occupations) Key Findings Indicator Value Note Occupations analysed 502 Complete CNO-11 (SEPE taxonomy) Workers represented 22.46 million EPA Q4 2025 (final data) Weighted mean exposure 3.7 / 10 Calibrated for Spain (5 factors + sub-components) High-exposure occupations (score ≥7) 47 occupations 2,752,961 jobs (12.3% of employment) [≥6.5: 15.4%] Wage-exposure index 249,143M EUR Employment × Salary × Score/10 — a weighted index, not a prediction Score range 1.0 – 9.0 17 unique values in 0.5 increments Salary range 12,985 – 79,282 EUR/year 128 unique estimated values Inter-model validation — holistic (r) 0.715 100 occupations, Gemini 2.5 Pro vs GPT-4o blind Inter-model validation — sub-components (r) 0.953 177 rescaled occupations, GPT 5.4 vs Gemini 3 Sub-component coverage (D/C/F/R) 502/502 (100%) 177 rescaled + 325 informative Salary validation (MAPE) 4.96% 16 INE EES 2023 groups, post-correction Employment validation (1-digit) ±0.00% EPA Q4 2025 exact (API Tempus table 65134) Employment validation (2-digit) ±0.02% max drift Hard constraint on 62 EPA subgroups Unique employment values 499 / 502 4-layer cascade (EPA→Census→SEPE) Employment confidence A 382 occupations Census + SEPE chain Employment confidence B 120 occupations SEPE-synthetic chain Registered unemployed (SEPE) 10.02M (Dec 2024) Cross-validation signal: r = 0.73 vs contracts EU AI Act classification 51 high-risk, 4 limited, 447 minimal Reclassified v21 per Annex III use-case Methodology Exposure scores were generated by Gemini 2.5 Pro (temperature 0.2, structured rubric prompts) following the methodological lineage of Brynjolfsson et al. (2018) and Eloundou et al. (2023), adapted to the Spanish CNO-11 taxonomy with structural calibration. Five Spain-specific calibration factors are applied: DESI digitalisation index — DESI 2023, 69.8 points (3rd in EU). Spain ranks 11th in enterprise digital technology integration; 21.1% of Spanish firms (10+ employees) use AI (INE TIC Q1 2025; Banco de España EBAE 2025: ~20%). Sector variation: Services 25.7%, Industry 17.5%, Construction 11.4%. Moderation factor: 0.80 (agriculture) to 0.95 (technology/banking). Services sector weight — 74% of GDP (vs 68% EU average); tourism 12.4% of GDP. Employment protection — 3rd strictest in OECD. Unfair dismissal severance: 33 days/year (max 24 monthly payments). Friction factor: 1–5% by sector. EU AI Act — Regulation (EU) 2024/1689 classifies AI systems by use-case context, not occupations. Annex III high-risk contexts map to 51 occupations (3.03M workers). Moderation factor: 2–8%. AESIA supervision — Spain is the first EU country with an operational national AI supervisory agency (A Coruña, Real Decreto 729/2023). Fines up to 35M EUR or 7% of global turnover. Combined calibration: 11–12% reduction from base LLM scores. Formula: score_adjusted = score_base × factor_DESI × (1 - regulatory_friction) × (1 - labour_friction) Sub-component rescoring (v22–v26): 177 occupations concentrated at scores 7.0 and 5.0 were rescored via 4-dimensional decomposition (D: Digitalisation, C: AI Capability, F: Physical barrier, R: Regulatory friction) by GPT 5.4 Thinking and Gemini 3 Thinking independently. Formula: score_raw = (D×C/10) × (1−F/20) × (1−R/20). An additional 325 occupations received informative sub-components (same protocol, no score change), completing D/C/F/R coverage to 502/502 (100%). Employment data: 4-layer cascade combining EPA Q4 2025 (1-digit and 2-digit hard constraints), Census 2021 (3-digit structural weights), and SEPE 2024 contracts (4-digit relative weights). Result: 499 unique employment values across 502 occupations. Salary data: EES 2023 (INE, table 28186) at 2-digit CNO, adjusted with educational premia and intra-group variance proxies from France (INSEE) and Portugal (INE-PT). MAPE: 4.96%. Unemployment data: SEPE registered unemployed (parados) at 4-digit CNO, December and June 2024, with gender breakdowns. Contracts-parados correlation: r = 0.73. Validation Stack Employment validation EPA Q4 2025 totals reproduced with ±0.00% at 1-digit (max difference: 47 persons over 22.46M) and ±0.02% at 2-digit (62 subgroups, hard constraint). AI score validation — holistic (inter-model) 100 stratified occupations blind-rescored by GPT-4o. Pearson r = 0.715, ICC(2,1) = 0.701, weighted κ = 0.667. Bland-Altman: no proportional bias, 95% limits [−2.77, +3.33]. AI score validation — sub-components (inter-model) 177 rescaled occupations scored by GPT 5.4 Thinking and Gemini 3 Thinking. Combined r = 0.953, 88% within ±1.0 point, Bland-Altman 95% limits [−1.12, +1.43]. 325 informative occupations: F dimension 91% agreement, overall r = 0.802 vs holistic score. Salary validation MAPE 4.96% against 16 INE EES 2023 groups. All deviations under 10%. Adversarial multi-model review Methodology subjected to adversarial review by 7+ independent AI models (Grok ×2, Perplexity ×2, Manus ×2, Gemini, Claude). 16 issues tracked; 14 fixed, 1 addressed with empirical anchoring, 1 documented limitation. Sensitivity Analysis Scenario Weighted mean High exposure (≥7) Wage-exposure index Current calibration (base) 3.7 / 10 12.3% 249,143M EUR All factors −20% ~4.8 / 10 Non-linear estimate ~321,600M EUR All factors +20% ~3.1 / 10 Non-linear estimate ~214,400M EUR Threshold sensitivity (±0.5): high exposure ranges from 11.9% (≥7.5) to 15.4% (≥6.5) — a 3.5pp band, narrowed from 13.2pp pre-rescoring. Adoption Context Scores measure theoretical AI capability, not current adoption. Three independent sources anchor the gap: INE TIC Q1 2025 (21.1% of firms use AI), Banco de España EBAE 2025 (~20%), Anthropic Economic Index (33% task-level in most-exposed occupations). Effective economy-wide task adoption: ~7%. The 5 calibration factors capture ~11–12% (structural friction); the remaining ~70pp reflects organisational barriers. Reader guidance: apply a 70–80% discount for current impact, varying by sector and firm size. Limitations Exposure scores are theoretical estimates, not predictions of job displacement. 4-digit employment figures are proportional estimates via a 4-layer cascade. Precise at 1-digit (±0.00%) and 2-digit (±0.02%); approximate at 4-digit. Calibration factors are expert judgement without empirical back-testing. Interaction effect quantified: max 0.06 points (negligible). France/Portugal salary proxies assume structural similarity. MCVL is an identified but unused validation source. Scores were generated in a single pass per model (holistic). Estimated intra-model reproducibility: ±0.5 points. Sub-component rescoring used two independent models. The analysis is static (March 2026 snapshot). No regional variation, part-time/full-time distinction, or job creation modelling. Salary clustering: 128 unique values; 445/502 occupations share a salary with at least one other. Root cause: INE publishes at 2-digit CNO only. Self-employed (~3.3M) are excluded from the salary survey by design. No temporal deflator on EES 2023 salaries (reference year 2022). Applying CPI would break the MAPE validation anchor. Interactive Dashboard The dashboard (single-page React application) provides four views: treemap (sector-level with drill-down), detailed treemap (occupation-level), scatter plot (salary vs exposure), and sortable list. Filters include sector, score range, and sort options. The detail panel shows the full profile including D/C/F/R sub-components for all 502 occupations, EU AI Act classification, impact typology, and wage-exposure sub-index. Hosted at https://alvarodenicolas.com/interactive/empleos-ia/index.html with a mirror at https://empleo-ai.netlify.app/. Comparative Positioning Dimension This analysis (v26B) Karpathy 'Jobs' (US) OECD AI Exposure ILO GenAI Index Taxonomy CNO-11 (502, Spain) ONET/SOC (~800, US) ~400 ISCO (cross-country) ISCO (cross-country) Scoring LLM + 5 calibration + 4D sub-components LLM single-pass (Gemini Flash) Expert + ONET tasks GPT-4 task scoring Inter-model validation r=0.715 holistic, r=0.953 sub-comp, Bland-Altman None published Expert panel None published Sub-component coverage 502/502 (100%) D/C/F/R None None None Regulatory mapping EU AI Act (3 risk levels, 51 high-risk) None None None Salary cross-reference Yes (128 values, MAPE 4.96%) Yes (BLS direct) No No Unemployment data Yes (SEPE parados + gender) No No No Adoption anchoring INE TIC 21.1% + BdE + Anthropic None None None Documentation 32pp PDF, 47 technical notes Code README Working paper Working paper US Comparative Analysis Parameter US (Karpathy) Spain (v26) Primary cause Mean exposure ~4.6 3.7 Physical services weight + 5-factor calibration + sub-components % high exposure (≥7) ~27% 12.3% [11.9%–15.4%] Smaller knowledge economy + sub-component rescoring Regulatory classification Not included 3 EU AI Act levels (51 high-risk) No US federal AI framework Labour friction factor Not applied 1–5% by sector OECD 3rd strictest employment protection Salary granularity ~800 direct values 128 adjusted values INE publishes at 2-digit level Employment granularity Direct per occupation 4-layer cascade: 499 unique values EPA anonymises CNO at 1–2 digit Sub-component rescoring Not included 502 occ × 4 sub-comp × 2 models v22–v26 Empirical adoption anchor Not included INE TIC 2025: 21.1% Harmonised Eurostat survey Changes from Previous Zenodo Version (v26/v13 → v26B/v13) This is a reproducibility-only update. No data, scores, or methodology changed. v26B updates: prompts_and_scripts_v23.zip → prompts_and_scripts_v26.zip (22 files → 42 files across 7 directories). The v23 archive contained raw sub-component data for only 177 of 502 occupations, making the "100% validation coverage" claim unverifiable from the published archive. The v26 archive adds all 325 informative sub-component raw outputs (11 GPT lots + 11 Gemini lots), the processing script (postprocess_325.py), patch JSON, full results CSV, divergence analysis, and model-specific prompts. Every occupation in the v13 dataset now traces to verifiable raw model outputs. metodologia_v26.pdf uploaded (replaces v23 PDF). 32 pages, 47 technical notes, 5 appendices. Primary dashboard URL updated to https://alvarodenicolas.com/interactive/empleos-ia/index.html (Netlify mirror retained). build_v26.py included in reproducibility package (ReportLab script generating the methodology PDF from v13 dataset). Previous changes (v23/v10 → v26B/v13): Dataset v10 → v13: Fields per record: 26 → 34 (+8 new fields) Employment: 4-layer cascade replacing Census-only redistribution. 202 → 499 unique values. EPA 2-digit drift reduced from ±540% to ±0.02%. Unemployment: SEPE parados (Dec/Jun 2024) with gender breakdowns added. Gender: Contract and unemployment splits by sex added. Employment confidence tiers (A/B) added. Sub-component D/C/F/R coverage: 177 → 502 occupations (100%). Divergence flag (flag_divergencia_gt2) added for 158 occupations. Weighted mean exposure: 3.74 → 3.66. High exposure: 12.5% → 12.3%. Wage index: 257,060M → 249,143M EUR. Methodology v23 → v26B: Pages: 22 → 32. Technical notes: 43 → 47. Section 4 rewritten for 4-layer employment cascade. New: parados/gender documentation, employment confidence tiers, cross-validation analysis. New: empirical adoption anchoring (INE TIC 2025 21.1%, BdE ~20%, Anthropic 33%) with interaction effect bound. New: 325 informative sub-components completing 502/502 D/C/F/R coverage. New notes [44]–[47]: calibration interaction bound, INE TIC anchoring, adoption-exposure gap magnitude, informative sub-components. Issue tracker: 12 fixed → 14 fixed, 1 addressed, 1 limitation (15/16 resolved). Technical Notes The methodology document (v26B) contains 47 technical notes grouped by topic: occupational inventory (notes 1–2), employment data (3–9), unemployment and gender (10–12), salary data (13–16), AI scoring (17–26), calibration and adoption (44–46), sub-components (47), EU AI Act (27–29), visualisation (30), limitations (31–36), and validation (37–43). Key additions since v23: [44] Interaction effect between calibration factors: max delta 0.06 points (negligible). [45] INE TIC Q1 2025 empirical anchor: 21.1% AI adoption (Services 25.7%, Industry 17.5%, Construction 11.4%). [46] Adoption-exposure gap: economy-wide task adoption ~7%; 70–80% discount guidance. [47] Informative sub-components for 325 occupations: F dimension 91% agreement, r = 0.802 vs holistic. Files spain_502_v13_subcomp_complete.json — Complete dataset (502 occupations, 34 fields, v13) metodologia_v26B.pdf — Full methodology (32 pages, 47 technical notes, 5 appendices) prompts_and_scripts_v26.zip — Reproducibility package (42 files across 7 directories: raw D/C/F/R scores for all 502 occupations [177 rescaled in JSON + 325 informative in markdown lots], scoring prompts, batch instructions, processing scripts, adversarial protocol, issue tracker, dataset versions v9/v10/v13, README) empleo-ia-standalone_v26B.html — Interactive standalone dashboard Keywords artificial intelligence, labour market, employment, Spain, AI exposure, occupational risk, EU AI Act, CNO-11, EPA, automation, interactive dashboard, inter-model validation, Bland-Altman, treemap, wage-exposure index, AESIA, sub-component scoring, unemployment, gender, SEPE, adoption gap, reproducibility License Creative Commons Attribution 4.0 International (CC BY 4.0) Language Spanish (dataset, justifications, dashboard UI); English (this description, methodology notes bilingual) Resource Type Dataset + Interactive Visualisation + Methodology Document Related Identifiers https://alvarodenicolas.com/interactive/empleos-ia/index.html (IsSupplementedBy — interactive dashboard, primary URL) https://empleo-ai.netlify.app/ (IsSupplementedBy — interactive dashboard, mirror) https://zenodo.org/records/19165098 (IsNewVersionOf — previous version v26/v13) Brynjolfsson, E., Mitchell, T. & Rock, D. (2018). "What Can Machines Learn, and What Does It Mean for Occupations and the Economy?" AEA Papers and Proceedings, 108:43-47. (References) Eloundou, T., Manning, S., Mishkin, P. & Rock, D. (2023). "GPTs are GPTs." OpenAI/UPenn. arXiv:2303.10130v5. (References) Frey, C. B. & Osborne, M. A. (2017). "The Future of Employment." Technological Forecasting and Social Change, 114:254-280. (References) Regulation (EU) 2024/1689 (EU AI Act). Annex III, Arts. 5 and 6. (References) Anthropic. The Anthropic Economic Index. February 2026. (References) INE. Encuesta sobre el uso de TIC y del comercio electrónico en las empresas. Q1 2025. (References) Banco de España. EBAE 2025. "La adopción de la inteligencia artificial en las empresas españolas." (References) Nedelkoska, L. & Quintini, G. (2018). "Automation, skills use and training." OECD Social, Employment and Migration Working Papers, No. 202. (References) Acemoglu, D. (2024). "The Simple Macroeconomics of AI." Economic Policy. (References)



