Saetov, I. (2026). Data and code for: The proximity of Turkic to other language groups
收藏资源简介:
Data, code and figures for a study measuring the structural proximity of the Turkic languages to a worldwide sample of language groups. The study asks whether a single number for proximity exists and finds that the answer depends on the level of language measured. Mongolic is the one group that stays at the top across levels; below it the rankings diverge. Contents data/ derived tables: the 53x195 Grambank feature matrix, per-level distance matrices and Turkic-proximity rankings for the six levels (grammar, usage, phonology, inventory, lexical semantics, neural), the spatiophylogenetic variance partition, per-level geographic controls, level-agreement and concordance tables, within-Turkic comparisons, sample coverage, Glottolog coordinates, and the neural nearest-neighbour table behind Figure 7. figures/ the nine article figures, labels in English. code/ scripts that rebuild the derived data from open sources, and make_figures.py, which rebuilds all nine figures from data/ alone (matplotlib + numpy). Levels of language Grammar from Grambank and the decorrelated GBI sets; usage from Universal Dependencies with the gradient word-order measures of Levshina; phonology and segment inventory from URIEL; lexical semantics from CLICS3 colexifications; neural representation from multilingual-e5-small sentence embeddings over FLORES-200. Morphology from UniMorph is used qualitatively only, dictionary coverage being too uneven for a reliable distance. Reproducing python code/make_figures.py rebuilds figures/ from data/ (no downloads). The data-building scripts incode/ fetch the raw sources directly; pin the release of each database, since results depend on it. The neural level depends on multilingual-e5-small with torch 2.2.2 and transformers 4.44.2. The Bayesian spatiophylogenetic model uses a fixed random seed. Raw sources are not redistributed; see SOURCES.md. Scope and limits The study measures similarity. Claims about language history lie outside what this measure supports. The variance partition describes the structure of similarity and does not identify a mechanism, so convergent typology, ancient contact and common descent remain equally compatible with the data. Licences Code under MIT (LICENSE-CODE). Derived data and figures under CC BY 4.0 (LICENSE-DATA). When reusing the data, cite the primary sources in SOURCES.md, which carry their own licences. Citation Saetov, I. (2026). Data and code for: The proximity of Turkic to other language groups (stable at the top, level-dependent below). Zenodo. https://doi.org/10.5281/zenodo.21613740



