Annotated database of Slovenian adverbs
收藏资源简介:
Annotated database of Slovenian adverbs Petra Mišmaš Marko Simonović Stefan Milosavljević This database presents a morphological annotation of Slovenian adverbs. It includes the 4,520 most frequent items annotated as adverbs in the Gigafida 2.0 corpus (deduplicated). The items were extracted from the corpus using the CQL query [tag="R.*"] on a random sample of 10,000,000 lines in the NoSketch Engine in May 2024. Only items occurring ten times or more were included. The initial list contained several items from other categories (primarily adjectives) as well as forms with typos. Such cases were marked as excluded but kept on the list for traceability. As a result, the database provides detailed morphological annotation for 3,904 adverbial items. Column-by-column overview We start by listing the columns in the database and indicating which property is annotated in each of them. Column A: ID All items in the database are annotated with consecutive numbers, and this column contains a unique number assigned to each item. Column B: Item This column lists the citation form (lemma) of each item. Column C: Frequency This column provides the frequency of each individual lemma. Column D: Included This column distinguishes between items we consider actual adverbs in the relevant sense from all other items. Words marked with 1 are included in the annotation, while items marked with a 0 are excluded. The reasons for exclusion are: the item not being an adverb and the item being misspelled. Columns E–K: Suffix 1 to Suffix 7 These columns list the specific suffixes contained in each adverb. Suffix 1 is the one closest to the root, followed by Suffix 2, and so on. The aim was to pursue maximal decomposition. Therefore, for instance, the adverb tako ‘so’ was decomposed into t-a-k-o, based on its relation to, e.g., k-a-k-o ‘how’, t-a-m ‘there’, t-a ‘this’. See Appendix for the specific decisions regarding the annotation. Column L: Ending By far most adverbs in Slovenian end in mid vowels and are homophonous with the neuter form of adjectives (e.g. biološko ‘biological-NEUT’ and ‘biological-ly’). In some cases there is a distinction in stress (e.g. lép-o ‘beautiful-NEUT’ vs. lepó ‘beautiful-ly’). Even in the latter case, the final vowel behaves as inflectional morphology, since it disappears in the comparative (e.g. lépše ‘more beautiful-ly’). A similar pattern is found with a small number of other adverbs, for instance bliz-u ‘nearby’ (cf. its comparative form bliž(j)-e). In the dataset, we coded an adverbial ending in Column L for: (i) adverbs ending in a vowel that have a clearly related adjective with forms lacking that vowel, and (ii) other adverbs whose final vowel disappears in the comparative. The value entered in this column is the vowel as realised in the surface form of the adverb, not a hypothesised underlying exponent. Thus, biološko ‘biologically’ is annotated as ending in -o, and tuje ‘in a foreign way’ as ending in -e, even though these are plausibly allomorphs of the same morpheme and their realisation is likely conditioned by the preceding consonant. In this context, it is relevant that many adverbs contain ‘fossilised’ case forms, e.g. ponoči ‘by night’ or povrhu ‘on top of that’. These final vowels are arguably still perceived as separate morphemes, but they are not inflectional, which is why we list them as suffixes. By the same token, the prepositions appearing in such adverbs are marked as prefixes. Column M: Compound base Adverbs whose base is a compound receive 1 in this column; all others receive 0. If an adverb is annotated with 1, only the right-hand component of the compound is decomposed and only for suffixes. For example, the first part of enakomerno ‘uniformly’ is en-ak containing the suffix -ak, but this is not annotated separately, whereas the second part is decomposed into mer-n-o. Loan adverbs are marked as having a compound base if the base components are also used independently in Slovenian in a compound-like structure. For instance, the base of tipološko ‘tipologically’ is tipolog, which contains tip (an independent word meaning ‘type’) and -log, also attested in words such as psiholog ‘psychologist’ and arheolog ‘archaeologist’. Column N: Prefixes If the adverb has a prefix, the prefix is listed in this column. If the adverb has several prefixes, these are listed in the column and are separated by a plus sign. The rightmost prefix is the one closest to the root/base. Items that could be taken to be a prefix, but an unprefixed version of the base (or a version with a different prefix) is not attested, are given in brackets. For instance zanikrno ‘sloppily’ has (za) in this column, since the annotator has the intuition that za is a prefix in this word, but *nikrn is not attested. If it was unclear whether an element constituted a single prefix or could be further decomposed, we provided a proposed decomposition. One such example is neizpodbitno ‘uncontentiously’, where the prefixes iz- and pod- also exist (as do the prepositions iz, pod, and izpod). For this reason, the form was annotated as having prefixes ne+iz+pod. Prefixes in loanwords are annotated in this column if the version without the prefix (or with some other prefix) also exists in Slovenian. E.g., iracionalno ‘irrationally’ is annotated as having the prefix i- because racionalno ‘rational’ also exists. On the other hand, re- in represivno ‘repressive’ is not given in this column, because *presivno does not exist in Slovenian. Appendix: Specific decisions for the annotation of suffixes in columns E–J The general criterion for annotating an element as a suffix was its occurrence in multiple words and/or in combination with other suffixes. Crucially, this means that we also attempted to decompose elements sometimes considered a single suffix. For example, -kast in siv-kast-o ‘grayishly’ was annotated as siv+k+ast-o, since both -k and -ast are independently attested suffixes (kič-ast-o ‘kitschy’, ljub-k-o ‘cute’). Especially in the domain of borrowed words, in some cases, it was impossible to reconstruct the underlying representation of suffixes that only appear before palatalising suffixes. For instance, in sarkastično ‘sarcastically’, the sequence -ič- can, in principle, be underlyingly -ik-, -ic-, or -ič-, as all these underlying representations could lead to the surface allomorph -ič-. In such cases, we opted for analogy with comparable words whose intermediate bases do surface as independent words. In this case, an analogy can be made with words like logistično ‘logistically’, with the base logistika ‘logistics’. As a consequence, sarkastično was annotated as sarkast+ik+n-o. Some nominal bases display so-called stem extensions, which occur throughout the paradigm of the noun (e.g. im-e ‘name’ has the genitive singular im-en-a, dative singular im-en-u etc.). Stem extensions like en were not annotated as derivational suffixes, so that e.g., po-imen-sk-o is annotated as having only the suffix sk. When -j is present in the declension of the noun, it was not annotated as a suffix. For example, loanwords ending on a vowel, e.g., kupe ‘compartment’ with the genitive singular kupeja, where the adverb kupejevsko ‘relatedly to a compartment’ is annotated as kupe+ov+sk+o. Phonologically conditioned allomorphs were generally annotated as a single morpheme: The morpheme -ov- systematically surfaces as -ev- after a set of consonants traditionally termed soft (j, c, č, ž, š). This morpheme was annotated as -ov- regardless of its surface form. E.g., kraljevsko ‘royally’ was annotated as kralj+ov+sk-o. This allows us to make a distinction between the morpheme -ov- and the morpheme -ev-, which can appear after all consonants (e.g. in u-po-št-ev-a-j-e ‘considering’). Consonant-initial suffixes can trigger the insertion of an epenthetic vowel in some forms. In such cases, the suffix was annotated in the version without the epenthetic vowel. For instance, bolezensko ‘pathologically’ was annotated as bol+ez+n+sk-o, because the e does not surface in the genitive singular of the base noun bolezen, bolezn-i etc. This allows us to make a distinction between -n and -en, which always surfaces with the vowel. An example with the latter suffix is zasluženo ‘deservedly’, annotated as za+služ+en-o. When annotating adverbs derived from passive participles, we assume that theme vowels are present whenever they can be reconstructed from the surface form. The clearest case are passive participles in -a-n, where the theme vowel is preserved. For example, za+dih+an-o ‘breathlessly’ from za-dih-a-ti se ‘to become breathless’, where the theme vowel a is annotated separately. We also annotated the theme vowel in cases where it survives in the form of a consonant (e.g., ne-s-pre-men-j-en-o ‘unchangedly’ from s-pre-men-i-ti ‘change’, annotated as ne+s+pre+men+i+en-o), multiple consonants (e.g., iz-gub-lj-en-o ‘in a lost way’ from iz-gub-i-ti ‘to lose’, annotated as iz+gub+i+en-o) or through the palatalisation of the preceding consonant (e.g., u-ska-j-en-o ‘in a coordinated way’ from u-sklad-i-ti ‘coordinate’, annotated as having the suffixes i+en since dj systematically palatalises to j). In the verbal domain, the suffix -ov systematically varies between two allomorphs: ov (e.g., in the infinitive pot-ov-a-ti ‘to travel’) and u (e.g. in pot-u-je-mo ‘we travel’). In derivation, the former allomorph is generally used, e.g., in pot-ov-a-l-en ‘related to travel’. However, in the so called active adjectival participles, as well as the adverbials derived from them, the latter allomorph surfaces, e.g., in pod-cen-j-uj-oč ‘underestimating’ and pod-cen-j-uj-oč-e ‘underestimatingly’. In both cases, we annotated the relevant morpheme as ov, so that pod-cen-j-uj-oč-e was annotated as pod+cen+i+ov+j+oč-e. If a derivational affix generally does not trigger the palatalisation of the preceding consonant, then in all cases where palatalisation does occur, the word is assumed to contain an additional palatalising morpheme -j-. For instance, the suffix -en generally does not trigger palatalisation (e.g., polst-en ‘made of felt’ derived from polst ‘felt’). Therefore, košč-en-o ‘as a bone’ derived from kost ‘bone’ is assumed to have an extra palatalising morpheme and was annotated as kost+j+en-o.
### 斯洛文尼亚语副词标注数据库 作者:Petra Mišmaš、Marko Simonović、Stefan Milosavljević 本数据库针对斯洛文尼亚语副词开展形态学标注。数据源自Gigafida 2.0语料库中频次最高的4520个被标注为副词的条目(已去重)。2024年5月,研究者通过NoSketch Engine平台,针对1000万行随机语料样本执行CQL查询`[tag="R.*"]`提取上述条目,仅保留出现频次≥10次的条目。 初始清单包含若干非副词类条目(主要为形容词)及存在拼写错误的形式。此类条目被标记为排除项,但仍保留于清单中以保证可追溯性。最终,本数据库为3904个副词条目提供了详细的形态学标注。 ## 逐列说明 我们将逐一介绍数据库的列项,并说明每一列所标注的属性。 ### 列A:ID 数据库中所有条目均被赋予连续编号,本列存储为每个条目分配的唯一编号。 ### 列B:Item 本列列出每个条目的引用形式(即词元lemma)。 ### 列C:Frequency 本列提供每个词元的出现频次。 ### 列D:Included 本列用于区分符合相关定义的真实副词与其他条目。标记为1的条目将被纳入标注集,标记为0的条目则被排除。排除原因包括: 1. 非副词条目 2. 存在拼写错误。 ### 列E至列K:后缀1至后缀7(Suffix 1 to Suffix 7) 本系列列列出每个副词所包含的具体后缀。后缀1为距离词干最近的后缀,后缀2次之,依此类推。 本标注遵循最大程度拆分的原则。例如,副词`tako`(意为“如此”)被拆分为`t-a-k-o`,这是基于其与`k-a-k-o`(意为“如何”)、`t-a-m`(意为“那里”)、`t-a`(意为“这个”)的词源关联。 具体标注决策详见附录。 ### 列L:词尾(Ending) 斯洛文尼亚语绝大多数副词以中元音结尾,且与形容词的中性形式同音异义(例如`biološko`,既可表示“生物学的(中性)”,也可表示“生物学地”)。部分情况下二者可通过重音区分(例如`lép-o`“漂亮的(中性)”与`lepó`“漂亮地”)。即便如此,该词尾仍属于屈折形态,因为在比较级形式中该元音会消失(例如`lépše`“更漂亮地”)。少量其他副词也遵循类似模式,例如`bliz-u`“附近”(其比较级形式为`bliž(j)-e`)。 本数据集在列L中为副词标注词尾的规则如下: (i) 以元音结尾,且存在无该元音的相关形容词形式的副词; (ii) 词尾元音在比较级中会消失的其他副词。 本列所填值为副词表层形式中的元音,而非假设的底层语素。例如,`biološko`“生物学地”被标注为以`-o`结尾,`tuje`“以异域方式”被标注为以`-e`结尾,尽管二者可能属于同一语素的同位异形体,且其实现形式可能受前置辅音制约。 在此语境下,需注意许多副词包含“石化”的格形式,例如`ponoči`“在夜间”或`povrhu`“除此之外”。此类词尾元音通常仍被视为独立语素,但不属于屈折形态,因此我们将其列为后缀。同理,此类副词中的介词成分会被标记为前缀。 ### 列M:复合词基底(Compound base) 若副词的基底为复合词,则本列标记为1;其余情况标记为0。 若某副词被标记为1,则仅对复合词的右侧成分进行拆分,且仅拆分其后缀部分。例如,副词`enakomerno`“统一地”的第一部分`en-ak`包含后缀`-ak`,但该部分不会被单独标注;而第二部分则被拆分为`mer-n-o`。 若外来副词的基底成分在斯洛文尼亚语中可独立使用,且可构成类复合词结构,则该副词被标记为拥有复合词基底。例如,`tipološko`“从类型学角度而言”的基底为`tipolog`,其中`tip`(意为“类型”,可独立成词)与`-log`(亦见于`psiholog`“心理学家”、`arheolog`“考古学家”等词)均为独立可证的成分。 ### 列N:前缀(Prefixes) 若副词包含前缀,则本列列出该前缀;若包含多个前缀,则以加号分隔各前缀,最右侧的前缀距离词干/基底最近。 对于看似前缀,但无该前缀的基底形式(或带有其他前缀的基底形式)未被证实的条目,其前缀会被置于括号中。例如,`zanikrno`“草率地”在本列中标注为`(za)`,因为标注者认为`za`在此词中为前缀,但`*nikrn`未被证实存在。 若无法明确某一成分是单个前缀还是可进一步拆分,则会给出推测的拆分方式。例如`neizpodbitno`“无争议地”,其中`iz-`与`pod-`均为独立存在的前缀(同时也作为介词`iz`、`pod`以及复合介词`izpod`使用),因此该词被标注为前缀序列`ne+iz+pod`。 若无该前缀(或带有其他前缀)的形式在斯洛文尼亚语中存在,则外来词中的前缀会被列入本列。例如`iracionalno`“非理性地”被标注为带有前缀`i-`,因为`racionalno`“理性地”是存在的。反之,`represivno`“压迫性的”中的`re-`则未被列入本列,因为`*presivno`在斯洛文尼亚语中不存在。 --- ## 附录:列E至列J的后缀标注具体决策 将某一成分标注为后缀的通用标准为:该成分可在多个词中出现,且/或可与其他后缀组合使用。至关重要的是,这意味着我们会尝试拆分有时被视为单个后缀的成分。例如,`siv-kast-o`“略带灰色的”中的`-kast`被拆分为`siv+k+ast-o`,因为`-k`与`-ast`均为独立可证的后缀(`kič-ast-o`“俗气的”、`ljub-k-o`“可爱的”)。 尤其在外来词领域,部分仅出现在腭化后缀之前的后缀,其底层表征难以重构。例如`sarkastično`“讽刺地”中的`-ič-`序列,理论上底层可表征为`-ik-`、`-ic-`或`-ič-`,因为这三种底层形式均可生成表层同位异形体`-ič-`。此类情况下,我们会参照那些中间基底可作为独立词出现的同类词。例如参照`logistično`“从逻辑角度而言”,其基底`logistika`“物流”可作为独立词使用。因此,`sarkastično`被标注为`sarkast+ik+n-o`。 部分名词基底存在所谓的词干延伸,该延伸贯穿名词的所有变格形式(例如`im-e`“名字”的属格单数为`im-en-a`,与格单数为`im-en-u`等)。此类词干延伸(如`en`)不会被标注为派生词缀,因此`po-imen-sk-o`仅被标注为包含后缀`sk`。 当名词变格中出现`-j`时,不会将其标注为后缀。例如以元音结尾的外来词`kupe`“隔间”,其属格单数为`kupeja`,对应的副词`kupejevsko`“与隔间相关地”被标注为`kupe+ov+sk+o`。 语音条件下的同位异形体通常被标注为单个语素: 1. 语素`-ov-`在一组传统称为软辅音(`j`、`c`、`č`、`ž`、`š`)之后会系统地实现为`-ev-`。无论表层形式如何,该语素均被标注为`-ov-`。例如`kraljevsko`“王室地”被标注为`kralj+ov+sk-o`。这一标注方式可区分语素`-ov-`与语素`-ev-`,后者可在所有辅音后出现(例如`u-po-št-ev-a-j-e`“考虑到”)。 2. 以辅音开头的后缀在部分形式中会触发插入嵌合元音。此类情况下,后缀会以无嵌合元音的形式标注。例如`bolezensko`“病理地”被标注为`bol+ez+n+sk-o`,因为在基底名词`bolezen`的属格单数`bolezn-i`等形式中不会出现`e`。这一标注方式可区分`-n`与`-en`,后者始终带有元音。后者后缀的示例为`zasluženo`“应得地”,被标注为`za+služ+en-o`。 在标注源自被动分词的副词时,我们假定只要可从表层形式重构出主题元音,则该元音即存在。最明确的案例为以`-a-n`结尾的被动分词,其中主题元音得以保留。例如`za+dih+an-o`“气喘吁吁地”源自`za-dih-a-ti se`“变得气喘吁吁”,其中主题元音`a`被单独标注。我们还会在以下情况中标注主题元音:主题元音以辅音形式留存(例如`ne-s-pre-men-j-en-o`“未改变地”源自`s-pre-men-i-ti`“改变”,被标注为`ne+s+pre+men+i+en-o`)、以多个辅音形式留存(例如`iz-gub-lj-en-o`“以丢失的方式”源自`iz-gub-i-ti`“丢失”,被标注为`iz+gub+i+en-o`)或通过前置辅音的腭化实现(例如`u-ska-j-en-o`“以协调的方式”源自`u-sklad-i-ti`“协调”,由于`dj`系统腭化为`j`,因此被标注为包含后缀`i+en`)。 在动词领域,后缀`-ov-`存在两种系统交替的同位异形体:`ov`(例如不定式`pot-ov-a-ti`“旅行”)与`u`(例如`pot-u-je-mo`“我们旅行”)。在派生过程中,通常使用前者的同位异形体,例如`pot-ov-a-l-en`“与旅行相关的”。然而在所谓的主动形容词分词以及源自此类分词的副词中,会出现后者的同位异形体,例如`pod-cen-j-uj-oč`“低估的”与`pod-cen-j-uj-oč-e`“低估地”。两种情况下,我们均将相关语素标注为`ov`,因此`pod-cen-j-uj-oč-e`被标注为`pod+cen+i+ov+j+oč-e`。 若某派生词缀通常不会触发前置辅音的腭化,则在所有发生腭化的情况下,该词均被假定包含额外的腭化语素`-j-`。例如后缀`-en`通常不会触发腭化(例如`polst-en`“毛毡制的”源自`polst`“毛毡”)。因此,源自`kost`“骨头”的`košč-en-o`“以骨头的方式”被假定包含额外的腭化语素,被标注为`kost+j+en-o`。



