Global database of vowel dissimilation
收藏资源简介:
This dataset consists of descriptive data about vowel dissimilation patterns across 116 languages and 133 unique patterns, representing an array of 38 linguistic families, including Afro-Asiatic, Austronesian, Indo-European, Mayan, Niger-Congo, Sino-Tibetan, Turkic, as well as isolates like Basque and Ainu. The largest representation within the language families is for the Austronesian family, represented with 28 patterns observed in 23 languages, followed by 17 patterns in 15 Indo-European languages, 10 patterns in 10 Atlantic-Congo and 8 Afro-Asiatic languages, respectively, and 8 distinct patterns found in 7 Mayan languages. Other smaller language families are also represented and make up half of the total patterns collected in the survey. Linguistic isolates are represented with 5 patterns in 3 languages. The dissimilation is signaled from the input-output pairs and the morpho-syntactic context in which the pattern is noticed (a specific morpheme, group of morphemes, or unlimited with respect to the morpho-syntax). The information and the examples are sourced from grammatical descriptions, primarily reference and descriptive grammars, but also dictionaries, wordlists, corpora, various online materials like forums, song lyrics, news portals, magazines, religious texts. Valuable sources are phonological descriptions and phonological studies on individual languages as well as descriptive and theoretical papers offering analyses of dissimilative patterns in various frameworks. Genetic information for individual languages is sourced from Glottolog v. 5.2, supplemented with information from the sources themselves, where necessary. For example, in some cases the language name in Glottolog is different from the name in the source, in which case priority is given to the source. Linguistics systems are presented in an alphabetical order according to the major name followed by the modifier. This means that variants of the same major linguistic system are presented after one another, like in the case of Basque, Guere etc. Every linguistic system / languoid is identified with an ISO-3 code or Glottocode (if the ISO-3 Code is not available, usually for smaller variants), language family, area(s) where spoken and the list of sources the data are retrieved from. The general information is followed by the marker `phonological' or `morpho-phonological' depending on the observed nature of the pattern. Next is the information on the dissimilative regularity and the morpho-syntactic context in which the pattern functions, followed by the data, represented as lists of examples showing the regular pattern in contrast to dissimilative, including notes about exceptions and general phonological tendencies in the language. The amount of data available is sadly not uniform and is in several cases scarce. In some cases only the representative examples are available and in some all of the available data are taken into account, even if that meant the pattern is represented with five examples. Columns in the dataset: Language Identification & Metadata glottocode - Unique language identifier from Glottolog v. 5.2 (e.g., adyg1241) language.x - Language name (e.g., "Adyghe") iso.x - ISO 639-3 code (e.g., ady) family - Language family (e.g., "Abkhaz-Adyge") subfamily - Subgroup (e.g., "Circasian") language_glottolog - Glottolog's standardized language name language_glottolog.1 - Secondary Glottolog reference iso.y - Alternate ISO code (if different from iso.x) level - Language/dialect classification ("language" or "dialect") Geographic Data area - Macro-region (e.g., "Eurasia", "Africa") latitude - Decimal degrees longitude - Decimal degrees countries - ISO country codes (e.g., "RU;TR") Dissimilation Patterns VD.type - Pattern type (P = phonological, MP = morpho-phonological) feature.INPUT - Underlying vowel feature (e.g. [+low]) feature.OUTPUT - Resulting feature (e.g. [-low]) feature.CONTEXT - Phonological context triggering change other.features - Additional relevant features (e.g. [+round]) type.of.identity - What kind of identity is necessary for dissimilation ("full" or "partial") vowel.length - Sensitivity to vowel length ("no", "feeds", "bleeds") adjacent - Locality condition ("syllable", "root node", "foot", "unlimited", "variable") Morphosyntactic Context morphemes.involved - Morpho-syntactic context (e.g., "pl", "poss") another - Secondary morpheme category (if applicable) class - Word class affected ("noun", "verb", "both") direction - "regressive" or "progressive" dissimilation trigger - From where dissimilative originate ("prefix", "suffix", "root") location - Locus of change ("root", "suffix", etc.) Additional Features prosody.related - Stress/tone involvement ("yes"/"no") alternative - Alternative value to dissimilative (e.g. "default", "harmony", "reduplication") feature_change - Descriptive string (e.g. [[+low]] → [[-low]]) morpheme_categories - Grammatical categories (e.g. "pers/num") Genealogical & Classification Data affiliation - Language family with sub-branches subclassification - Detailed genealogical tree countries - Repeat of ISO country codes Example Entry kase1253 Kasem xsm Atlantic-Congo Grusi MP [-low] [+low] [+high] [+round] partial feeds syllable pl no noun regressive suffix root no default Kasem Kasem xsm language Africa 11.0824 -1.39076 BF;GH Atlantic-Congo, Volta-Congo, North Volta-Congo, Gur, Central Gur, Southern Central Gur, Grusi, Northern Grusi, Nuna-Kasem (East_Kasem:1,Fere:1,Lela:1,Nuclear_Kasem:1,Nunuma:1,West_Kasem:1)kase1253:1; [[-low]] → [[+low]] pl
本数据集涵盖116种语言的133种独特元音异化(vowel dissimilation)模式描述性数据,涉及38个语系,包括亚非语系、南岛语系、印欧语系、玛雅语系、尼日尔-刚果语系、汉藏语系、突厥语系,以及巴斯克语、阿伊努语等孤立语言。语系收录占比最高的为南岛语系,涵盖23种语言中的28种异化模式;其次为印欧语系,涵盖15种语言中的17种模式;大西洋-刚果语系与亚非语系分别涵盖10种语言中的10种模式、7种语言中的8种模式;此外玛雅语系涵盖7种语言中的8种独特模式。其余小型语系亦有收录,占本次调研收集模式总数的一半。孤立语言则涵盖3种语言中的5种异化模式。 本次收录的异化现象通过输入-输出对以及该模式被观测到的形态句法(morpho-syntactic)语境(特定语素(morpheme)、语素组,或无形态句法限制)进行标注。本数据集的信息与示例主要来自语法描写文献(以参考语法与描写性语法为主),同时涵盖词典、词表、语料库,以及论坛、歌词、新闻门户、杂志、宗教文本等各类在线资源。具有较高学术价值的数据源包括针对单一语言的音系描写与音系研究,以及各类分析框架下对异化模式进行探讨的描写性与理论性论文。 单一语言的系属信息来自格洛托洛格(Glottolog)v.5.2版本,必要时辅以数据源本身的信息。例如,部分情况下格洛托洛格中的语言名称与数据源中的名称存在差异,此时优先采用数据源的名称。本数据集按语言体系的主名称加修饰语的字母顺序排列,即同一主语言体系的变体将依次排列,如巴斯克语、格雷语等案例。 每个语言体系/语群均通过ISO 639-3代码或格洛托科德(Glottocode,若ISO 639-3代码不可用,通常用于小型变体)、语系、使用区域以及数据来源列表进行标识。基础信息之后会标注“音系(phonological)”或“形态音系(morpho-phonological)”,具体取决于该模式的观测属性。随后是异化规则以及该模式运行的形态句法语境信息,最后是数据部分:以示例列表形式呈现常规模式与异化模式的对比,同时包含该语言的例外情况与通用音系倾向说明。遗憾的是,可用数据量并不统一,部分案例的数据较为稀缺。部分情况下仅能获取代表性示例,部分案例则会纳入所有可用数据,即便该模式仅能以5个示例进行呈现。 数据集列项如下: ### 语言标识与元数据(Language Identification & Metadata) - glottocode:来自格洛托洛格(Glottolog)v.5.2的唯一语言标识符(示例:adyg1241) - language.x:语言名称(示例:"Adyghe") - iso.x:ISO 639-3代码(示例:ady) - family:语系(示例:"Abkhaz-Adyge") - subfamily:语族分支(示例:"Circasian") - language_glottolog:格洛托洛格标准化语言名称 - language_glottolog.1:格洛托洛格二级参考名称 - iso.y:备用ISO代码(若与iso.x不同) - level:语言/方言分类("language"或"dialect",即“语言”或“方言”) ### 地理数据(Geographic Data) - area:宏观区域(示例:"Eurasia", "Africa",即“欧亚大陆”“非洲”) - latitude:纬度(十进制度) - longitude:经度(十进制度) - countries:ISO国家代码(示例:"RU;TR") ### 异化模式(Dissimilation Patterns) - VD.type:模式类型(P = 音系型,MP = 形态音系型) - feature.INPUT:底层元音特征(例如[+low]) - feature.OUTPUT:结果特征(例如[-low]) - feature.CONTEXT:触发音系变化的语境 - other.features:其他相关特征(例如[+round]) - type.of.identity:异化所需的同一性类型("full"或"partial",即“完全”或“部分”) - vowel.length:对元音长度的敏感度("no", "feeds", "bleeds",即“无”“馈合”“渗溢”,为音系学常用表述) - adjacent:邻接条件("syllable", "root node", "foot", "unlimited", "variable",即“音节”“词根节点”“音步”“无限制”“可变”) ### 形态句法语境(Morphosyntactic Context) - morphemes.involved:形态句法语境(示例:"pl", "poss",即“复数”“所属格”) - another:二级语素类别(如适用) - class:受影响的词类("noun", "verb", "both",即“名词”“动词”“两者皆是”) - direction:"regressive"或"progressive" dissimilation,即“逆异化”或“顺异化” - trigger:异化现象的起源位置("prefix", "suffix", "root",即“前缀”“后缀”“词根”) - location:变化发生的位置("root", "suffix"等) ### 附加特征(Additional Features) - prosody.related:重音/声调参与情况("yes"/"no",即“是”/“否”) - alternative:异化替代值(示例:"default", "harmony", "reduplication",即“默认值”“和谐”“重叠”) - feature_change:特征变化描述字符串(示例:[[+low]] → [[-low]]) - morpheme_categories:语法范畴(示例:"pers/num",即“人称/数”) ### 系属与分类数据(Genealogical & Classification Data) - affiliation:含分支的语系归属 - subclassification:详细系属树形结构 - countries:重复的ISO国家代码 示例条目: kase1253 Kasem xsm Atlantic-Congo Grusi MP [-low] [+low] [+high] [+round] partial feeds syllable pl no noun regressive suffix root no default Kasem Kasem xsm language Africa 11.0824 -1.39076 BF;GH Atlantic-Congo, Volta-Congo, North Volta-Congo, Gur, Central Gur, Southern Central Gur, Grusi, Northern Grusi, Nuna-Kasem (East_Kasem:1,Fere:1,Lela:1,Nuclear_Kasem:1,Nunuma:1,West_Kasem:1)kase1253:1; [[-low]] → [[+low]] pl



