Annotated database of nominalization pairs in Serbian
收藏资源简介:
Annotated database of nominalization pairs in Serbian The dataset contains annotations of (arguably) deverbal nominalizations in Serbian, based on 4,132 concordance lines extracted from the Serbian web corpus CLASSLA-web.sr (Ljubešić et al. 2024). The nominalizations in question include 60 native -nje nominalizations and 20 Latinate -cija nominalizations. The nominalizations were selected to form pairs of lemmas of two types: 20 -nje nominalizations derived from imperfective verbs were paired with 20 -nje nominalizations derived from their perfective counterparts; 20 -cija nominalizations were paired with 20 -nje counterparts. The database is part of a study aimed at determining the extent to which eventive deverbal nominalizations in English and Serbian exhibit verbal properties. The study focused on two key indicators of verbal properties: the presence of a genitive complement and the presence of plural forms of nominalisations. -nje nominalizations are the most productive type of deverbal nominalizations in Serbian, derived by adding the suffix -je to the passive participle form of the corresponding verb (1). They can be derived both from imperfective (e.g. rešavati ‘solve.ipfv’ → rešavanje) and perfective (rešiti ‘solve.pfv’ → rešenje) verbs. Imperfective-derived and perfective-derived nominalizations differ structurally, syntactically, prosodically and semantically. Whereas nominalizations derived from imperfective verbs are fully productive, semantically transparent and exhibit prosodic faithfulness, perfective-derived nominalizations have limited productivity, lexicalized semantics and they are prosodically unfaithful (Simonović & Arsenijević 2014; Kovačević 2021). (1) slikati → slikan + -je → slikanje paint.ɪɴꜰ paint.pass.ptcp paint.nominalization -cija nominalizations are Latinate nominalizations that primarily denote an event (e.g. manipulacija ‘manipulation’). They are very frequent, relatively productive and differ both in the stem and in the prosodic pattern from the related adjectives and verbs (Simonović & Arsenijević 2018). Selecting the pairs and testing the ‘process’ reading The selection process began with over 20 pairs of nominalizations per category, chosen ad hoc by the first author, with the additional criterion that each nominalization had at least 1,000 corpus hits. The categories comprised: (a) pairs consisting of imperfective-derived nominalizations in -nje and their corresponding perfective-derived counterparts, and (b) pairs consisting of -cija nominalizations and their -nje counterparts. A table was created in Google Spreadsheets, containing information on lemma frequency and frequency per million words, extracted from the corpus through a dedicated search of each nominalization. For example, the entry for the nominalization eliminisanje ‘elimination’ shows a lemma frequency of 7,608 and a frequency per million of 2.64. To ensure that each included nominalization can be interpreted as a process, a Corpus Query Language (CQL) search was performed using the pattern [lemma="proces"][lemma="x"], where x represents the specific nominalization. Pairs that did not yield any corpus hits were excluded from further analysis. Note that while the application of the “process” test meant some level of semantic control, the semantics of the nominalizations was not the primary criterion during the data selection process. Therefore, many included nominalizations can also have non-process meanings as well. For example, the word varijacija ‘variation’, related to the verb varirati ‘to vary’, generally can refer to the process of variation of a melodically and rhythmically structured entity (e.g., through changes in rhythm, tempo, etc.), but is far more common in its non-eventive meaning. The annotation in the database does not directly target the semantics of the nominalization, but rather focuses on its morpho-syntactic correlates, specifically the presence of a genitive complement and the ability to form a plural. Since, after the application of the process test, more than 20 pairs were still available per category, the pairs had to be ranked to select the most appropriate ones. Ranking the pairs of nominalizations The final selection was based on the lemma frequency ratios of the nominalizations within each pair. For example, among the pairs consisting of a -nje nominalization and the corresponding -cija nominalization, there was the pair eliminisanje ‘eliminating’ vs. eliminacija ‘elimination’. The lemma frequency of eliminisanje (7,608) was divided by the lemma frequency of eliminacija (20,202), yielding a ratio of 0.38. Within each category, the 20 pairs with ratios closest to 1 were selected to ensure the most balanced pairs in terms of frequency. Extracting concordances from the corpus The final dataset contains 50 valid example sentences per nominalization. Sentences found in titles and subtitles, as well as duplicate sentences, were excluded due to their different argument structures. Consequently, more than 50 concordances were initially extracted for each nominalization using dedicated CQL queries to ensure that 50 valid examples remained for analysis. Annotation process and criteria The concordances were annotated for several linguistic features, including whether the nominalization is used in a sentence, whether it is pluralized, the presence of a genitive complement, the presence of a possessive, the presence of a clausal complement, and the presence of a free relative clause as its complement. The presence of each structure is marked with 1, and its absence is marked with 0. Column-by-column explanation Column A (nominalization_type) indicates the type of nominalization in question, combining the derivational suffix and the aspect specification. All -nje nominalizations that have pairs in -cija turned out to be derived from bisapectual verbs, so that these nominalizations got the value -nje.BIASP. -nje.IPFV: imperfective-derived -nje nominalization (e.g. poništavanje ‘annulment’ derived from poništavati ‘annul.IPFV’); -nje.PFV: perfective-derived -nje nominalization (e.g. poništenje ‘annulment’ derived from poništiti ‘annul.PFV’ ); -nje.BIASP: biaspectual-derived -nje nominalization (e.g. aktiviranje ‘activating’); -cija - nominalizations in -cija (e.g. aktivacija ‘activation’). Column B (root) contains the root shared by the nominalizations within the pair (e.g. aktiv- in the pair aktiviranje/aktivacija). Column C (nominalization) contains the nominalization in question (e.g. aktiviranje). Column D (concordance) contains the concordance in which the particular nominalization appears. E.g., (2) is a concordance that contains the nominalization aktiviranje ‘activation’. (2) Aktiviranjem ovog člana pokreću se dvogodišnji pregovori [...] activation.INS.SG this article.GEN.SG trigger SE two-year negotiations ‘The activation of this article triggers two years of negotiations [...]’ Column E (sentence) indicates whether the context in which the nominalization in question is used constitutes a full sentence. Titles and subtitles were not considered as full sentences. For example, (3) received the value 0, as it represents a subtitle. (3) AKTIVIRANJE STARE IDEJE. activation old.GEN.SG idea.GEN.SG ‘THE ACTIVATION OF AN OLD IDEA.’ On the other hand, (4) received the value 1. (4) Aktiviranje naloga dovodi do ugovora između učesnika i LOVEORTHODOX. activation account.GEN.SG leads to contract between participants and LOVEORTHODOX ‘The activation of the account leads to the contract between the participants and LOVEORTHODOX.’ The value 0 was also assigned to sentences that were repetitions of already annotated sentences. Column F (plural) indicates whether the nominalization in question is in the plural. For example, aktivacijama ‘activations’ in (5) received the value 1. (5) [...] a svi posetioci će biti u prilici da uživaju u fantastičnom celodnevnom and all visitors will be in opportunity to enjoy in fantastic all-day zabavnom programu, specijalnim popustima, iznenađenjima i zanimljivim aktivacijama. entertainment program special discounts surprises and exciting activities.LOC.PL ‘[...] and all visitors will have the opportunity to enjoy a fantastic all-day entertainment program, special discounts, surprises, and exciting activities.’ Column G (genitive) indicates whether the right-hand context of the nominalization contains a genitive-marked complement that can be interpreted as an argument of the underlying verb. Both external and internal genitive arguments are taken into account. For instance, aktivaciju in (6) received the value 1. (6) Saradnja bi obuhvatala sva tri nivoa studija prava collaboration would encompass all three levels study.GEN.PL law.GEN.SG kao i aktivaciju već postojećeg instituta.gen.sg „gostujući profesor“. as well as activation.acc.sg already existing institute “visiting professor” ‘The collaboration would encompass all three levels of law studies, as well as the activation of the already existing "visiting professor" institute.’ Sometimes a token yields two possible interpretations, as in (7). (7) neracionalno zaduživanje građana irrational indebtedness/borrowing.NOM.SG citizen.GEN.PL ‘irrational indebtedness of the citizens’ / ‘irrational citizen’s borrowing’ Here, either citizens take on debt themselves, or someone or something, such as the state, indebts them. Therefore, a decision was made to include both possibilities under this factor — both internal and external genitive arguments — assigning them both the value 1. Column H (possessive) indicates the presence or absence of a possessive pronoun, or a possessive adjective ending in -ov and -in that can be interpreted as an argument of the nominalization in question. For example, (8) received the value 1. (8) [...] a njihovo aktiviranje je stvar vremena, a ne principa. and their activation.NOM.SG is matter time.GEN.SG and not principle.GEN.SG ‘[...] and their activation is a matter of time, not principle.’ Column I (clausal_complement) indicates the presence or absence of a finite clause introduced by the complementizer da as a complement of the nominalization in question. For example, (9) received the value 1. (9) [...] čija su zaduženja da promovišu kulturne manifestacije… whose are duties to promote cultural manifestations ‘[...] whose duties are to promote cultural manifestations…’ Column J (free_relative_complement) indicates the presence or absence of a free relative clause as a complement of the nominalization in question. For example, (8) received the value 1. (8) [...] pošto postoje ograničenja koliko često mogu da budu pozvani. because exist limitations how often can.3PL DA be.3SG invited ‘[...] since there are limitations on how often they can be invited.’ References Ljubešić, N, Rupnik, P. and Kuzman, T. (2024). Serbian web corpus CLASSLA-web.sr 1.0, Slovenian language resource repository CLARIN.SI, ISSN 2820-4042, http://hdl.handle.net/11356/1931 Simonović, M. and Arsenijević B. (2014). Regular and honorary membership: on two kinds of deverbal nouns in Serbo-Croatian. Lingue e Linguaggio, 2014 (2). 185-210. Simonović, M. and Arsenijević B. (2018). The importance of not belonging: Paradigmaticity and loan nominalizations in Serbo-Croatian. Open Linguistics, vol. 4, no. 1, 2018, pp. 418-437. https://doi.org/10.1515/opli-2018-0021 Kovačević, P. (2021). On the internal structure of Serbian -(n)je nominalizations. Acta Linguistica Academica, 68(4), 426–453. https://doi.org/10.1556/2062.2021.00434



