Quantifiers in Erzya and Moksha
收藏资源简介:
The dataset comprises quantifiers in the Erzya and Moksha languages, collected and analyzed from the MokshEr corpus. It includes a frequency list of quantifiers along with their inflected forms and an in-depth analysis of 2400 quantifiers. A breakdown of the analyzed forms is presented in Table 1. The frequency lists contain details about the selected quantifier stems: ламо* and аламо* in the Erzya subcorpus, and лам* and крж* in the Moksha subcorpus. Irrelevant forms, such as inflected verb forms ламолгадомс 'to become more', кржалгодомс 'to become less', were manually excluded from the frequency lists. After filtering the data, new percentages were calculated for items with at least 100 occurrences. Based on the frequency data, specific quantifiers were selected for detailed contextual analysis. In Erzya, the focus was on ламо 'many, much', ламо-т [many-PL] 'many, much', ламо-т-не [many-PL-DEF] 'many, much' (including case-marked forms), and аламо 'not many, few, little'. In Moksha, the corresponding forms analyzed were лама 'many, much', ламо-ц [many-POSS.3SG] 'many, much', ламо-т-не [many-PL-DEF] 'many, much' (including case-marked forms), and кржа 'little, few'. Table 1. The forms of quantifiers analysed in the dataset. Erzya ламо 300 ламот 300 ламотне (+ case endings) 300 аламо 300 Moksha лама 300 ламоц 300 ламотне (+ case endings) 300 кржа 300 The quantifiers were examined in context, with 10 words before and after the search item, and additional context included where necessary. The following features were considered in the analysis: Function of the quantifier Case marking of the noun Declension type of the noun Number marking on the noun Noun type (countable, uncountable, pronominal) Sentence type / phrase structure Grammatical function of the NP Word order Only quantifiers used in a relevant context were included. For instance, ламоц [many-POSS.3SG], typically used as a modifier directly preceding a definite plural noun, was not analyzed for word order or sentence type, as these features do not influence case and number marking of the noun. The function of the quantifier has two different abbreviations for adnominal function: adncore and adnobl. The former, adncore, includes NPs functioning as the subject, indefinite direct object, time adverbials, or expression of distance. In these grammatical functions, the noun is expected to be in the nominative, but the ablative marking can be influenced by the presence of the quantifier. Postpositions with a noun in the nominative are encoded as adnobl, as no case variation is expected in this function. Further abbreviations used for the function of quantifiers are pred for quantifiers in the predicate position, quant for so-called split structures that are comparable to the predicate function of the quantifier, and adv where the quantifier does not modify a noun. Quantifiers in adverbial function were not analyzed further. Case marking of the noun includes abbreviations from traditional grammatical description of the Mordvinic languages for cases and postpositions used in quantifier constructions (i.e., the postpositions e-/ez-, E jutko, M jotka). Nouns are further analyzed for three declension types: basic, definite, and possessive. Number marking distinguishes singular and plural, though in the possessive declension singular and plural number cannot always be distinguished. Nouns in the basic declension were treated as singular forms, though number distinction is possible only in the nominative. The noun type includes information about whether the noun is countable or uncountable, or belongs to the pronominal word class. Preliminary analyses were also conducted on sentence type, grammatical function of the NP, and word order, though these features require further, more systematic research. The dataset contains numerous non-verbal clauses, which must be carefully reviewed in follow-up studies. In retrospect, documenting the coninuity of the quanifier construction would have been more informative than surface word order alone. For instance, quantifiers and nouns that follow each other in the surface word order may still belong to different NPs. Intervening parts of speech between quantifier and noun were not marked; for example, intervening place or time adverbials were not annotated, despite their potential impact on NP structure. Abbreviations 3 = 3rd person, DEF = definite declension, PL = plural, POSS = possessive declension, SG = singular Reference to the corpus: MokshEr = MokshEr -corpus v3.0 (Moksha, Erzya). (https://finno-ugric-corpora.utu.fi/cqpweb/).



