Annotated database of Slovenian adverbs
收藏资源简介:
Annotated database of Slovenian adverbs Petra Mišmaš Marko Simonović Stefan Milosavljević This database presents a morphological annotation of Slovenian adverbs. It includes the 4,520 most frequent items annotated as adverbs in the Gigafida 2.0 corpus (deduplicated). The items were extracted from the corpus using the CQL query [tag="R.*"] on a random sample of 10,000,000 lines in the NoSketch Engine in May 2024. Only items occurring ten times or more were included. The initial list contained several items from other categories (primarily adjectives) as well as forms with typos. Such cases were marked as excluded but kept on the list for traceability. As a result, the database provides detailed morphological annotation for 3,904 adverbial items. Column-by-column overview We start by listing the columns in the database and indicating which property is annotated in each of them. Column A: ID All items in the database are annotated with consecutive numbers, and this column contains a unique number assigned to each item. Column B: Item This column lists the citation form (lemma) of each item. Column C: Frequency This column provides the frequency of each individual lemma. Column D: Included This column distinguishes between items we consider actual adverbs in the relevant sense from all other items. Words marked with 1 are included in the annotation, while items marked with a 0 are excluded. The reasons for exclusion are: the item not being an adverb and the item being misspelled. Columns E–K: Suffix 1 to Suffix 7 These columns list the specific suffixes contained in each adverb. Suffix 1 is the one closest to the root, followed by Suffix 2, and so on. The aim was to pursue maximal decomposition. Therefore, for instance, the adverb tako ‘so’ was decomposed into t-a-k-o, based on its relation to, e.g., k-a-k-o ‘how’, t-a-m ‘there’, t-a ‘this’. See Appendix for the specific decisions regarding the annotation. Column L: Ending By far most adverbs in Slovenian end in mid vowels and are homophonous with the neuter form of adjectives (e.g. biološko ‘biological-NEUT’ and ‘biological-ly’). In some cases there is a distinction in stress (e.g. lép-o ‘beautiful-NEUT’ vs. lepó ‘beautiful-ly’). Even in the latter case, the final vowel behaves as inflectional morphology, since it disappears in the comparative (e.g. lépše ‘more beautiful-ly’). A similar pattern is found with a small number of other adverbs, for instance bliz-u ‘nearby’ (cf. its comparative form bliž(j)-e). In the dataset, we coded an adverbial ending in Column L for: (i) adverbs ending in a vowel that have a clearly related adjective with forms lacking that vowel, and (ii) other adverbs whose final vowel disappears in the comparative. The value entered in this column is the vowel as realised in the surface form of the adverb, not a hypothesised underlying exponent. Thus, biološko ‘biologically’ is annotated as ending in -o, and tuje ‘in a foreign way’ as ending in -e, even though these are plausibly allomorphs of the same morpheme and their realisation is likely conditioned by the preceding consonant. In this context, it is relevant that many adverbs contain ‘fossilised’ case forms, e.g. ponoči ‘by night’ or povrhu ‘on top of that’. These final vowels are arguably still perceived as separate morphemes, but they are not inflectional, which is why we list them as suffixes. By the same token, the prepositions appearing in such adverbs are marked as prefixes. Column M: Compound base Adverbs whose base is a compound receive 1 in this column; all others receive 0. If an adverb is annotated with 1, only the right-hand component of the compound is decomposed and only for suffixes. For example, the first part of enakomerno ‘uniformly’ is en-ak containing the suffix -ak, but this is not annotated separately, whereas the second part is decomposed into mer-n-o. Loan adverbs are marked as having a compound base if the base components are also used independently in Slovenian in a compound-like structure. For instance, the base of tipološko ‘tipologically’ is tipolog, which contains tip (an independent word meaning ‘type’) and -log, also attested in words such as psiholog ‘psychologist’ and arheolog ‘archaeologist’. Column N: Prefixes If the adverb has a prefix, the prefix is listed in this column. If the adverb has several prefixes, these are listed in the column and are separated by a plus sign. The rightmost prefix is the one closest to the root/base. Items that could be taken to be a prefix, but an unprefixed version of the base (or a version with a different prefix) is not attested, are given in brackets. For instance zanikrno ‘sloppily’ has (za) in this column, since the annotator has the intuition that za is a prefix in this word, but *nikrn is not attested. If it was unclear whether an element constituted a single prefix or could be further decomposed, we provided a proposed decomposition. One such example is neizpodbitno ‘uncontentiously’, where the prefixes iz- and pod- also exist (as do the prepositions iz, pod, and izpod). For this reason, the form was annotated as having prefixes ne+iz+pod. Prefixes in loanwords are annotated in this column if the version without the prefix (or with some other prefix) also exists in Slovenian. E.g., iracionalno ‘irrationally’ is annotated as having the prefix i- because racionalno ‘rational’ also exists. On the other hand, re- in represivno ‘repressive’ is not given in this column, because *presivno does not exist in Slovenian. Appendix: Specific decisions for the annotation of suffixes in columns E–J The general criterion for annotating an element as a suffix was its occurrence in multiple words and/or in combination with other suffixes. Crucially, this means that we also attempted to decompose elements sometimes considered a single suffix. For example, -kast in siv-kast-o ‘grayishly’ was annotated as siv+k+ast-o, since both -k and -ast are independently attested suffixes (kič-ast-o ‘kitschy’, ljub-k-o ‘cute’). Especially in the domain of borrowed words, in some cases, it was impossible to reconstruct the underlying representation of suffixes that only appear before palatalising suffixes. For instance, in sarkastično ‘sarcastically’, the sequence -ič- can, in principle, be underlyingly -ik-, -ic-, or -ič-, as all these underlying representations could lead to the surface allomorph -ič-. In such cases, we opted for analogy with comparable words whose intermediate bases do surface as independent words. In this case, an analogy can be made with words like logistično ‘logistically’, with the base logistika ‘logistics’. As a consequence, sarkastično was annotated as sarkast+ik+n-o. Some nominal bases display so-called stem extensions, which occur throughout the paradigm of the noun (e.g. im-e ‘name’ has the genitive singular im-en-a, dative singular im-en-u etc.). Stem extensions like en were not annotated as derivational suffixes, so that e.g., po-imen-sk-o is annotated as having only the suffix sk. When -j is present in the declension of the noun, it was not annotated as a suffix. For example, loanwords ending on a vowel, e.g., kupe ‘compartment’ with the genitive singular kupeja, where the adverb kupejevsko ‘relatedly to a compartment’ is annotated as kupe+ov+sk+o. Phonologically conditioned allomorphs were generally annotated as a single morpheme: The morpheme -ov- systematically surfaces as -ev- after a set of consonants traditionally termed soft (j, c, č, ž, š). This morpheme was annotated as -ov- regardless of its surface form. E.g., kraljevsko ‘royally’ was annotated as kralj+ov+sk-o. This allows us to make a distinction between the morpheme -ov- and the morpheme -ev-, which can appear after all consonants (e.g. in u-po-št-ev-a-j-e ‘considering’). Consonant-initial suffixes can trigger the insertion of an epenthetic vowel in some forms. In such cases, the suffix was annotated in the version without the epenthetic vowel. For instance, bolezensko ‘pathologically’ was annotated as bol+ez+n+sk-o, because the e does not surface in the genitive singular of the base noun bolezen, bolezn-i etc. This allows us to make a distinction between -n and -en, which always surfaces with the vowel. An example with the latter suffix is zasluženo ‘deservedly’, annotated as za+služ+en-o. When annotating adverbs derived from passive participles, we assume that theme vowels are present whenever they can be reconstructed from the surface form. The clearest case are passive participles in -a-n, where the theme vowel is preserved. For example, za+dih+an-o ‘breathlessly’ from za-dih-a-ti se ‘to become breathless’, where the theme vowel a is annotated separately. We also annotated the theme vowel in cases where it survives in the form of a consonant (e.g., ne-s-pre-men-j-en-o ‘unchangedly’ from s-pre-men-i-ti ‘change’, annotated as ne+s+pre+men+i+en-o), multiple consonants (e.g., iz-gub-lj-en-o ‘in a lost way’ from iz-gub-i-ti ‘to lose’, annotated as iz+gub+i+en-o) or through the palatalisation of the preceding consonant (e.g., u-ska-j-en-o ‘in a coordinated way’ from u-sklad-i-ti ‘coordinate’, annotated as having the suffixes i+en since dj systematically palatalises to j). In the verbal domain, the suffix -ov systematically varies between two allomorphs: ov (e.g., in the infinitive pot-ov-a-ti ‘to travel’) and u (e.g. in pot-u-je-mo ‘we travel’). In derivation, the former allomorph is generally used, e.g., in pot-ov-a-l-en ‘related to travel’. However, in the so called active adjectival participles, as well as the adverbials derived from them, the latter allomorph surfaces, e.g., in pod-cen-j-uj-oč ‘underestimating’ and pod-cen-j-uj-oč-e ‘underestimatingly’. In both cases, we annotated the relevant morpheme as ov, so that pod-cen-j-uj-oč-e was annotated as pod+cen+i+ov+j+oč-e. If a derivational affix generally does not trigger the palatalisation of the preceding consonant, then in all cases where palatalisation does occur, the word is assumed to contain an additional palatalising morpheme -j-. For instance, the suffix -en generally does not trigger palatalisation (e.g., polst-en ‘made of felt’ derived from polst ‘felt’). Therefore, košč-en-o ‘as a bone’ derived from kost ‘bone’ is assumed to have an extra palatalising morpheme and was annotated as kost+j+en-o.



