遇见数据集

Annotated database of deverbal nominalisations in Timok Torlak

收藏
Zenodo2026-05-28 更新2026-05-29 收录
官方服务:

资源简介:

Annotated database of deverbal nominalisations in Timok Torlak Anđela Babić¹, Marko Simonović², Mirjana Mirić¹¹ Institute for Balkan Studies, Serbian Academy of Sciences and Arts² University of Graz This dataset contains corpus-derived and manually annotated information on deverbal nominalisations in the Timok variety of Torlak, extracted from the Spoken Torlak Dialect Corpus 1.0 (Vuković 2020). The dataset was created to investigate morphological, aspectual, and syntactic properties of nominalisations in the classes -nje/-će and -cija. The nominalisations were extracted from the Spoken Torlak Dialect Corpus 1.0 (Vuković 2020), which was accessed at the following link: https://www.clarin.si/noske/run.cgi/corp_info?corpname=torlak As noted in Vuković (2020), the corpus consists of semi-orthographic transcripts of 86.5 hours of recordings from locations evenly distributed across the Timok area of the Torlak dialect zone; see Vuković (2021) for a more detailed description of the corpus. Procedure for obtaining the data Our search for nominalisation tokens made use of the corpus feature that allows speaker type to be selected. We included all speakers from the Timok area and excluded passages marked as having been uttered by the researchers. In order to extract all target tokens, we constructed regular expressions covering all inflectional forms of each target nominalisation. Based on our preliminary searches, we established that, as in standard BCMS, certain generalisations hold for the segments preceding -nje and -će in nominalisations, which allowed us to restrict our searches. Specifically, -nje is always preceded by the segments a or e, whereas -će is always preceded by the segments e, i, or u. On the basis of this observation, and drawing on our knowledge of Torlak nominal morphology, we formulated the regular expressions in (1). (1) Regular expressions for Torlak forms on -nje, -će and -cija nominals [word=".*(a|e|é|á)nj(a|e|é|á|ata|eto|éto|áta)"] [word=".*(e|é|i|í|ú|u)(ć|č|q|ḱ)(e|é|á|ata|eto|éto|áta)"] [word=".*(c|s|z)(i|í)j(a|ata|á|áta|e|ete|é|éte|u|utu|ú|útu)"] Initial inspection of the data showed that some standard BCMS case forms were attested in passages which otherwise display typical Torlak features. We therefore conducted additional searches, for the sake of completeness, to target genitive, dative/locative and instrumental case forms that our initial queries did not capture. This led us to introduce additional regular expressions. (2) [word=".*(a|e|é|á)nj(u|em|ima)"] [word=".*(e|é|i|í|ú|u)(ć|č|q|ḱ)(u|em|ima)"] [word=".*(c|s|z)(i|í)j(i|om|ama)"] All extracted tokens were entered into a Google Sheets table and assigned a unique ID. The first five columns of the dataset comprise the ID and the columns inherited from the corpus search output: corpus reference, left context, KWIC, and right context. First five columns of the dataset Column Explanation ID The numeric identifier of each token in the dataset. Reference The reference in the Spoken Torlak Dialect Corpus 1.0 (Vuković 2020). Left Left context. KWIC KWIC. Right Right context. Establishing inclusion criteria Column: Deverbal & target variety After compiling the initial dataset, we screened all tokens against two inclusion criteria: that they were forms of an -nje/-će or -cija nominalisation; and that they had a related verb. Tokens that did not meet the two inclusion criteria were coded as 0 in the column Deverbal & target variety. Tokens that met both criteria but either exhibited case forms not characteristic of Timok Torlak — specifically genitive, dative/locative, or instrumental forms — were coded as 2. All remaining tokens, that is, deverbal nominalisations of the target classes occurring in Torlak forms, were coded as 1. The released dataset contains all 2,979 extracted tokens, also including items coded as 0. Of these, 606 tokens were assigned the value 1 and 23 tokens the value 2. The 606 target tokens were subsequently annotated for the remaining properties of the nominalisations in question. Annotating features of the nominalisations For all tokens coded as 1 in the column Deverbal & target variety, we annotated a set of additional features concerning the morphological class of the nominalisation, its normalised lemma, its relation to standard BCMS, and the aspectual value of the related verb. The column nje/će distinguishes between the two nominalisation classes targeted in the dataset. Tokens belonging to the -nje/-će class were coded as 1, as in krštenje ‘christening’, tkanje ‘weaving’, mlevenje ‘grinding’, and raspeće ‘crucifixion’. Nominalisations in -cija were coded as 0, as in operacija ‘operation’, orijentacija ‘orientation’, okupacija ‘occupation’, and izolacija ‘isolation’. The column lemma gives the normalised citation form of the nominalisation. This makes it possible to group together orthographic, prosodic and inflectional variants of the same item, as well as the forms featuring different case and number markers, as well as the postpositive definite article. For example, the attested forms tkanjé, tkanje, and tkánja are assigned the lemma tkanje, while forms such as édenje, edénje, and jédenje are normalised as jedenje. Similarly, raspéḱe is normalised as raspeće, and operácija as operacija, while pričánje and pričánjeto were assigned to the lemma pričanje. The column standard_match records whether the token corresponds to an existing form in standard BCMS, disregarding prosodic differences. Forms such as krštenje ‘baptism/baptising’, tkanje ‘weaving’, operacija ‘operation’, pijenje ‘drinking’, and venčanje ‘wedding’ were coded as 1, since they correspond to standard BCMS forms. Forms such as práenja ‘making’, rékanje ‘saying/telling’, tepánje ‘beating’, slikuvánje ‘photographing’, and prostuvánje ‘forgiving’ were coded as 0, since they do not correspond straightforwardly to standard BCMS forms. The column base_verb contains the related verb from which the nominalisation was analysed as derived. The verb is given in the feminine past participle form, for example krstila ‘baptised’ for krštenje ‘baptism/baptising’, tkala ‘wove’ for tkanje ‘weaving’, operisala ‘operated’ for operacija ‘operation’, mlela ‘ground/milled’ for mlevenje ‘grinding/milling’, raspela ‘crucified’ for raspeće ‘crucifixion’, and vaskrsnula ‘resurrected’ for vaskrsenje ‘resurrection’. This format was chosen because it provides a stable citation form that also makes aspectual and stem-related information transparent, while also taking into account the fact that Torlak lacks the infinitive. Finally, the column PFV encodes the aspectual value of the related verb. Perfective verbs were coded as 1, as in raspela ‘crucified’ for raspeće ‘crucifixion’, oprostila ‘forgave’ for oproštenje ‘forgiveness/forgiving’, vaskrsnula ‘resurrected’ for vaskrsenje ‘resurrection’, rekla ‘said’ for rekanje ‘saying/telling’, and oslobodila ‘liberated’ for oslobođenje ‘liberation’. Imperfective and biaspectual verbs were coded as 0, as in tkala ‘wove’ for tkanje ‘weaving’, operisala ‘operated’ for operacija ‘operation’, mlela ‘ground’ for mlevenje ‘grinding’, krstila ‘baptised’ for krštenje ‘baptism/baptising’, and čitala ‘read/was reading’ for čitanje ‘reading’. Annotating the token and its complements The remaining columns annotate properties of the token in its immediate syntactic context. These columns record whether the nominalisation occurs with an overt complement and whether the nominalisation token is plural. The column gen_compl indicates whether the nominalisation is accompanied by a genitive-marked complement. Tokens with a genitive complement were coded as 1, as in zapalénje plúḱa ‘inflammation of the lungs, pneumonia’, proširénje srca ‘enlargement of the heart’, vaskrsénje mrtvih ‘resurrection of the dead’, and oprošténje gréhova ‘forgiveness of sins’. Tokens without such a complement were coded as 0. The column acc_compl records the presence of an accusative-marked complement without a preposition. Tokens with an accusative complement were coded as 1, while tokens without an accusative complement were coded as 0. For example, tkanjé ćilími ‘weaving carpets’ and operácija slépo crévo ‘operation on the appendix’ were marked as cases with an accusative complement. The column na_compl indicates whether the nominalisation occurs with a complement introduced by na. Tokens with a na-PP complement were coded as 1, as in kršténje na décu ‘baptising children’. Tokens without a na-complement were coded as 0. Finally, the column pl records whether the nominalisation is in a plural form. Crucially, for the purposes of this column, paucal forms following low numbers and some quantifiers were counted as plural. Plural tokens were coded as 1, while singular tokens were coded as 0. For example, bájanja ‘acts of incantation / charm-healing’ and pričánja ‘acts of telling / talking’ were coded as plural forms, whereas singular forms such as kršténje ‘christening’, tkanjé ‘weaving’, operácija ‘operation’, and raspéḱe ‘crucifixion’ were coded as 0. Remaining annotated columns of the dataset Column Explanation Deverbal & target variety Value 1: the token belongs to one of the target classes (-nje/-će or -cija nominals), has a related verb, and forms part of a Torlak utterance.Value 2: the token belongs to one of the target classes (-nje/-će or -cija), has a related verb, but is used in a case form that is not typical of Timok Torlak.Value 0: all other items. nje/će Value 1: -nje/-će nominals.Value 0: -cija nominals. lemma The citation form of the nominalisation in question, normalised. standard_match Value 1: the token corresponds to an existing form in standard BCMS, disregarding prosody.Value 0: the token does not correspond to an existing form in standard BCMS. base_verb The related verb in ᴘsᴛ.ᴘᴛᴄᴘ.ꜰ form. PFV Value 1: the related verb is perfective.Value 0: the related verb is imperfective or biaspectual. gen_compl Value 1: a genitive-marked complement is present.Value 0: no genitive complement is present. acc_compl Value 1: a bare accusative-marked complement is present.Value 0: no bare accusative-marked complement is present. na_compl Value 1: a na-PP complement is present.Value 0: no na-PP complement is present. pl Value 1: the nominalisation is in the plural.Value 0: the nominalisation is not in the plural. Acknowledgements This database was produced with financial support from the project What’s in a verb? Mapping Serbian verbs borrowed into Romani (Institute for Balkan Studies, Serbian Academy of Sciences and Arts (SASA), and Department of Slavic Studies, University of Graz), financed by the Ministry of Science, Technological Development and Innovation of the Republic of Serbia, in cooperation with Austria’s Agency for Education and Internationalisation (OeAD), within the programme of scientific and technological cooperation between the Republic of Serbia and the Republic of Austria for the period 2024–2026. Work on this database was also supported by the Austrian Science Fund (FWF), project Multifunctionality in morphology (Grant DOI: 10.55776/I6258; PI: Marko Simonović; University of Graz). References Vuković, Teodora. 2020. Spoken Torlak Dialect Corpus 1.0 (transcription). Slovenian language resource repository CLARIN.SI. ISSN 2820-4042. http://hdl.handle.net/11356/1281 Vuković, Teodora. 2021. Representing variation in a spoken corpus of an endangered dialect: The case of Torlak. Language Resources and Evaluation 55(3): 731–756. https://doi.org/10.1007/s10579-020-09522-4

提供机构:
Zenodo
创建时间:
2026-05-28
二维码
社区交流群
二维码
科研交流群
商业服务