Undoing Babel: AI, English, and the New Linguistic Infrastructure of Global Law
收藏资源简介:
"Undoing Babel: AI, English, and the New Linguistic Infrastructure of Global Law" This dataset accompanies the paper Undoing Babel: AI, English, and the New Linguistic Infrastructure of Global Law, which investigates the relationship between English language proficiency, colonial linguistic heritage, and a country’s readiness for AI governance. The core finding is that English proficiency—instrumented using colonial linguistic history—significantly predicts a country’s score on the 2024 Government AI Readiness Index (GAIRI). This suggests that English has become a global infrastructural language underpinning digital governance capacity. The econometric strategy uses Two-Stage Least Squares (2SLS) and Generalized Method of Moments (GMM-IV) estimation via the linearmodels Python package. Colonial language variables are used as instruments for English proficiency to address potential endogeneity. The Hansen J-test confirms instrument validity (p = 0.21). The analysis is fully reproducible and all Python scripts, datasets, and regression outputs are included. Files included: 2sls.py: Main estimation script (2SLS & GMM-IV models). ai.py, lgdp.py, orig.py: Supporting scripts for interaction effects and variable prep. README.md: Detailed project overview, variable definitions, and methodological notes. EF_EPI_2024_Ranking_with_Puerto_Rico.xlsx: English Proficiency Index data. 2024-GAIRI-data.xlsx: Government AI Readiness Index data. GDP_2023.xlsx: National GDP data (normalized and lagged). EEFR_All_States_and_Puerto_Rico.xlsx: U.S. state-level data (not used in global regressions). Sample size: N = 98 countriesSoftware: Python (pandas, statsmodels, linearmodels)License: Creative Commons Attribution 4.0 International (CC BY 4.0)DOI: 10.5281/zenodo.15635672



