Quantifiers in Erzya and Moksha
收藏资源简介:
The dataset comprises quantifiers in the Erzya and Moksha languages, collected and analyzed from the MokshEr corpus. It includes a frequency list of quantifiers along with their inflected forms and an in-depth analysis of 2400 quantifiers. A breakdown of the analyzed forms is presented in Table 1. The frequency lists contain details about the selected quantifier stems: ламо* and аламо* in the Erzya subcorpus, and лам* and крж* in the Moksha subcorpus. Irrelevant forms, such as inflected verb forms ламолгадомс 'to become more', кржалгодомс 'to become less', were manually excluded from the frequency lists. After filtering the data, new percentages were calculated for items with at least 100 occurrences. Based on the frequency data, specific quantifiers were selected for detailed contextual analysis. In Erzya, the focus was on ламо 'many, much', ламо-т [many-PL] 'many, much', ламо-т-не [many-PL-DEF] 'many, much' (including case-marked forms), and аламо 'not many, few, little'. In Moksha, the corresponding forms analyzed were лама 'many, much', ламо-ц [many-POSS.3SG] 'many, much', ламо-т-не [many-PL-DEF] 'many, much' (including case-marked forms), and кржа 'little, few'. Table 1. The forms of quantifiers analysed in the dataset. Erzya ламо 300 ламот 300 ламотне (+ case endings) 300 аламо 300 Moksha лама 300 ламоц 300 ламотне (+ case endings) 300 кржа 300 The quantifiers were examined in context, with 10 words before and after the search item, and additional context included where necessary. The following features were considered in the analysis: Function of the quantifier Case marking of the noun Declension type of the noun Number marking on the noun Noun type (countable, uncountable, pronominal) Sentence type / phrase structure Grammatical function of the NP Word order Only quantifiers used in a relevant context were included. For instance, ламоц [many-POSS.3SG], typically used as a modifier directly preceding a definite plural noun, was not analyzed for word order or sentence type, as these features do not influence case and number marking of the noun. The function of the quantifier has two different abbreviations for adnominal function: adncore and adnobl. The former, adncore, includes NPs functioning as the subject, indefinite direct object, time adverbials, or expression of distance. In these grammatical functions, the noun is expected to be in the nominative, but the ablative marking can be influenced by the presence of the quantifier. Postpositions with a noun in the nominative are encoded as adnobl, as no case variation is expected in this function. Further abbreviations used for the function of quantifiers are pred for quantifiers in the predicate position, quant for so-called split structures that are comparable to the predicate function of the quantifier, and adv where the quantifier does not modify a noun. Quantifiers in adverbial function were not analyzed further. Case marking of the noun includes abbreviations from traditional grammatical description of the Mordvinic languages for cases and postpositions used in quantifier constructions (i.e., the postpositions e-/ez-, E jutko, M jotka). Nouns are further analyzed for three declension types: basic, definite, and possessive. Number marking distinguishes singular and plural, though in the possessive declension singular and plural number cannot always be distinguished. Nouns in the basic declension were treated as singular forms, though number distinction is possible only in the nominative. The noun type includes information about whether the noun is countable or uncountable, or belongs to the pronominal word class. Preliminary analyses were also conducted on sentence type, grammatical function of the NP, and word order, though these features require further, more systematic research. The dataset contains numerous non-verbal clauses, which must be carefully reviewed in follow-up studies. In retrospect, documenting the coninuity of the quanifier construction would have been more informative than surface word order alone. For instance, quantifiers and nouns that follow each other in the surface word order may still belong to different NPs. Intervening parts of speech between quantifier and noun were not marked; for example, intervening place or time adverbials were not annotated, despite their potential impact on NP structure. Abbreviations 3 = 3rd person, DEF = definite declension, PL = plural, POSS = possessive declension, SG = singular Reference to the corpus: MokshEr = MokshEr -corpus v3.0 (Moksha, Erzya). (https://finno-ugric-corpora.utu.fi/cqpweb/).
本数据集收录了埃尔齐亚语(Erzya)与莫克沙语(Moksha)中的量化词(quantifier),数据源自莫克沙埃尔语料库(MokshEr corpus)并经整理分析。数据集包含量化词的频率列表及其屈折形式,并对2400个量化词展开了深度分析。表1展示了所分析形式的分类明细。 该频率列表收录了所选量化词词干的相关细节:埃尔齐亚子语料库中的ламо*和аламо*,以及莫克沙子语料库中的лам*和крж*。无关形式,如屈折动词形式ламолгадомс‘变得更多’、кржалгідомс‘变得更少’,已手动从频率列表中剔除。完成数据过滤后,对出现频次不少于100次的条目重新计算了占比。 基于频率数据,我们遴选了特定量化词开展详细的语境分析。在埃尔齐亚语中,重点考察的量化词包括:ламо“许多、大量”、ламо-т [many-PL]“许多、大量”、ламо-т-не [many-PL-DEF]“许多、大量”(含变格形式),以及аламо“不多、少量、少许”。在莫克沙语中,对应的分析形式为:лама“许多、大量”、ламо-ц [many-POSS.3SG]“许多、大量”、ламо-т-не [many-PL-DEF]“许多、大量”(含变格形式),以及кржа“少量、少许”。 表1 本数据集所分析的量化词形式 Erzya ламо 300 ламот 300 ламотне (+ case endings) 300 аламо 300 Moksha лама 300 ламоц 300 ламотне (+ case endings) 300 кржа 300 本次分析针对量化词开展语境考察,记录了目标词前后各10个词的上下文,并在必要时补充额外语境。分析过程中考量了以下特征: - 量化词的句法功能 - 名词的格标记 - 名词的变格类型 - 名词的数标记 - 名词类型(可数、不可数、代词类) - 句子类型/短语结构 - 名词短语(NP)的语法功能 - 词序 仅收录处于相关语境中的量化词。例如,通常直接修饰定指复数名词的ламоц [many-POSS.3SG],无需分析词序或句子类型,因为这些特征不会影响名词的格与数标记。 量化词的句法功能有两种用于定语功能的缩写:adncore与adnobl。其中adncore涵盖充当主语、不定直接宾语、时间状语或表距离的名词短语(NP)。在这类语法功能中,名词通常使用主格,但夺格标记可能受量化词影响。带主格名词的后置介词结构编码为adnobl,该功能下无变格变化。其他量化词功能缩写包括:pred(谓语位置的量化词)、quant(与量化词谓语功能类似的分裂结构)、adv(量化词未修饰名词的情况)。状语功能的量化词无需进一步分析。 名词的格标记涵盖莫尔达夫语(Mordvinic languages)传统语法描述中用于量化词结构的格与后置介词(即后置介词e-/ez-,埃尔齐亚语jutko,莫克沙语jotko)。我们进一步将名词划分为三种变格类型:基础变格、定指变格与领属变格。数标记区分单数与复数,不过在领属变格中无法始终区分单复数。基础变格的名词视为单数形式,尽管仅主格可实现数的区分。名词类型则指明该名词为可数、不可数,或属于代词词类。 我们还针对句子类型、名词短语语法功能与词序开展了初步分析,但这些特征需要更系统的后续研究。本数据集包含大量非谓语从句,需在后续研究中予以细致梳理。回顾来看,记录量化词结构的连续性或许比仅考察表层词序更具研究价值。例如,表层词序上相邻的量化词与名词可能分属不同的名词短语。量化词与名词之间的插入性词类未做标记;例如,插入的地点或时间状语即便可能影响名词短语结构,也未被标注。 缩写说明 3 = 第三人称,DEF = 定指变格,PL = 复数,POSS = 领属变格,SG = 单数 语料库来源: MokshEr = MokshEr语料库v3.0(莫克沙语、埃尔齐亚语)。(https://finno-ugric-corpora.utu.fi/cqpweb/)



