A Central Asian Language Survey
收藏资源简介:
TablesWe have documented language varieties (either Turkic or Indo-European) spoken in 23 test sites by 88 informants belonging to the major ethnic groups of Kyrgyzstan, Tajikistan and Uzbekistan (Karakalpaks, Kazakhs, Kyrgyz, Tajiks, Uzbeks, Yaghnobis). The recorded linguistic material concerns 176 words of the extended Swadesh list and will be made publically available with the publication of this paper. Phonological diversity is measured by the Levenshtein distance and displayed as a consensus bootstrap tree and as multidimensional scaling plots. Linguistic contact is measured as the number of borrowings, from one linguistic family into the other, according to a precision/recall analysis further validated by expert judgment. Concerning Turkic languages, the results of our sample do not support Kazakh and Karakalpak as distinct languages and indicate the existence of several separate Karakalpak varieties. Kyrgyz and Uzbek, on the other hand, appear quite homogeneous. Among the Indo-Iranian languages, the distinction between Tajik and Yaghnobi varieties is very clear-cut. More generally, the degree of borrowing is higher than average where language families are in contact in one of the many sorts of situations characterizing Central Asia: frequent bilingualism, shifting political boundaries, ethnic groups living outside the “mother” country.
本研究的相关表格记录了分属突厥语族与印欧语系的多种语言变体,这些语言分布于吉尔吉斯斯坦、塔吉克斯坦与乌兹别克斯坦境内的23个调研点,受访对象共88名,均来自当地主要族群:卡拉卡尔帕克人、哈萨克人、吉尔吉斯人、塔吉克人、乌兹别克人以及亚格诺比人。本次采集的语料包含扩充版斯瓦迪士词表(Swadesh list)中的176个词汇,本论文发表后将公开这批语料。语音多样性通过莱文斯坦距离(Levenshtein distance)进行量化,并以一致性自举树(consensus bootstrap tree)与多维标度图(multidimensional scaling plots)的形式呈现。语言接触程度以跨语系借词的数量进行量化,分析基于精确率-召回率分析(precision/recall analysis),并通过专家评估进一步验证。针对突厥语族语言,本次抽样结果不支持哈萨克语与卡拉卡尔帕克语为独立语言的观点,并表明存在多种不同的卡拉卡尔帕克语变体。反观吉尔吉斯语与乌兹别克语,则呈现出较强的同质性。在印伊语支(Indo-Iranian languages)语言中,塔吉克语变体与亚格诺比语变体的边界十分清晰。总体而言,在中亚典型的各类语言接触场景(如频繁的双语现象、变动的政治边界、族群旅居母国境外)中,跨语系借词比例均高于平均水平。




