Serbian loanverbs in Gurbet Romani (Knjaževac)
收藏资源简介:
Serbian loanverbs in Gurbet Romani (Knjaževac) Mirjana Mirić, Svetlana Ćirković, Marko Simonović The dataset comprises Serbian loanverbs attested in the Gurbet Romani variety spoken in Knjaževac (eastern Serbia). The loanverbs are extracted from transcripts of spoken narratives recorded from two groups of speakers: (a) adult native speakers (published in Ćirković & Mirić 2017 and Sikimić 2018), and (b) elementary-school students (unpublished, described in Mirić & Ćirković 2022). All speakers are bilingual in Gurbet Romani and Serbian. Data extraction and annotation procedure Each row in the dataset represents a single token of a loanverb. The tokens, along with their sentence contexts, were extracted from the transcripts by one of the researchers (Mirić). Two researchers (Mirić and Ćirković) independently annotated the tokens in the transcripts. The annotation was reviewed and cross-checked multiple times by both annotators and subsequently revised independently by a third researcher (Simonović). Summary of the columns The dataset contains 24 columns (A - X). A: Token number Contains the unique number of each token. B: Token Contains the token as used in the sentence from the transcript provided in the column T. C: Romani citation form Contains the lemma in Romani for each token, presented in the PRS.3SG form (due to the absence of the infinitive). D: Lemma (simplified) Contains the lemma in Romani, where verbs with the reflexive particle pe are unified with corresponding verbs without this particle. E.g., the intransitive verb vozil pe ‘to drive around’ and its transitive correspondent vozil ‘to drive/ride’ are unified as a single lemma vozil (pe). E: English translation Contains the English translation for each lemma. F: Included Indicates whether the token is included in the analysis: 0 = no, 1 = yes. Our analysis focused on the relationship between the BCMS verbs and the corresponding Romani verbs. The five tokens marked with 0 were excluded either because the corresponding BCMS verb could not be reliably reconstructed, or because the form of the Romani verb was exceptional compared to the rest of the dataset — making it plausibly a mistake. G: Romani characteristic vowel Contains the so-called ‘characteristic vowel’ in the Romani verb. This is the vowel preceding the ending -l in the citation form and has the possible values “i” or “o”, illustrated by volil ‘love’ and živol ‘live’, respectively. H: Romani characteristic vowel (binarised) Contains the numerically binarised version of column G (Romani characteristic vowel), where verbs with the characteristic vowel “i” are assigned the value 1, and those with the characteristic vowel “o” are assigned the value 0. I: Adaptation marker If the token contains a so-called adaptation marker (“sar” or “salj”), the exponent of the adaptation marker is provided in this column. If there is no adaptation marker, the value 0 is assigned. For instance, the token vozisaljam ‘we drove around’ has the value “salj” in this column, the token vozisardem ‘I drove’ has the value “sar”, whereas voziv ‘I drive’ has the value 0. J: Adaptation marker (binarised) Contains the binarised version of column I (Adaptation marked). The verbs which have an overt adaptation marker (“sar” or “salj”) were assigned the value 1, whereas those without an adaptation marker got the value 0. K: Verbal form Contains the information on the specific verbal form of the token. In this column we used the abbreviations from the Leipzig Glossing Rules. “PRS” = Present Tense, “PRF” = Perfect, “IMP”= Imperative, “IMPF” = Imperfect, “PTCP” = Deverbal participle. L: Verbal form (binarised) Contains a binarised version of column K (Verbal form) whereby all perfect forms have the value 1, whereas all other forms have the value 0. M: Base BCMS verb Contains the infinitive form of the verb in standard Bosnian/Croatian/Montenegrin/Serbian that the Romani verb is arguably based on. N: Base BCMS verb (certain) Indicates whether the base verb could be identified with certainty: 1 = yes, 0 = no. The value 0 is assigned in cases where: A corresponding BCMS verb could not be identified (e.g., brćil 'stir'); Multiple BCMS verbs could plausibly serve as the base (e.g., spremol 'prepare', which may derive from both the perfective verb spremiti and the imperfective spremati). O: Atypical integration pattern This column is used to separate those verbs which have an atypical integration pattern. The typical pattern is the one whereby the Romani characteristic vowel (“i” or “o”) is in the position where the theme vowel is in the BCMS verb, e.g. živeti → živol ‘live’, drogirati (se) → drogiril (pe) ‘use drugs’. All verbs that follow this pattern have the value 0 in this column. The atypical patterns, for which the value 0 is assigned, include: verbs where the BCMS theme vowel is maintained as part of the base to which the Romani characteristic vowel is added, e.g., upoznati → upoznail ‘meet’. verbs where a considerable sequence of the BCMS verb is deleted, including the portion preceding the BCMS theme vowel, e.g., reciklirati → recikol ‘recycle very few verbs where in the position of the characteristic vowel there is a vowel other than “i” or “o”, e.g., razumeti → razumel ‘understand’. P: Suffixless base Contains the information whether the base verb in BCMS contains a derivational suffix or not. All verbs which lack a derivational suffix have the value 1, whereas those with a derivational suffix have the value 0. The decision as to what counts as a derivational affix is necessarily somewhat arbitrary. For the purposes of this column, we only counted as derivational affixes those elements that are never part of the theme vowel material. For this reason, we did not consider the suffix n-, which precedes the theme vowel u/e, as a derivational affix, since arguably the same element is counted as a theme vowel in the TV class 0/ne. On the other hand, the suffix k- in sec-k-a-ti ‘cut’ is counted as a derivational affix, as it is never part of a theme vowel. Q: Base verb TV class Contains the information on the theme vowel class of the BCMS base verb. For this column we used the classification from the WeSoSlav database (Arsenijević et al. 2024). The available classes are: a/a, a/e, a/i, a/je, a/ne, e/e, e/i, i/i, u/e, 0/e, 0/ne. The names of the classes correspond to the exponents of the theme vowels in the infinitive and the PRS.3SG. For instance, the verb pištati ‘to whistle’ has the PRS.3SG form pišti and is therefore assigned to the TV class a/i. If multiple BCMS verbs, or a single verb that can belong to multiple TV classes, can serve as the base for the Romani verb, all relevant TV classes are listed, separated by “;”. For the single Romani verb for which we were unable to identify a corresponding BCMS verb, the value 0 was entered in this column. R: Preceding vowel Contains the nucleus of the syllable which precedes the characteristic vowel “i” or ”o” in the Gurbet Romani token. Note that apart from the vowels (a, e, i, o, u), “r” can also serve as a syllable nucleus in both Gurbet Romani and BCMS. S: Preceding vowel (binarised) Contains the binarised version of column R (Preceding vowel). If the preceding vowel is “i”, the value 1 was assigned, whereas in all other cases 0 was assigned. T: Romani sentence Contains the sentence from the transcript in Gurbet Romani from which the loanword token is taken. U: BCMS translation (sentence) Contains the Serbian translation of the example sentence in Gurbet Romani. This is applicable only to examples published in Sikimić (2018), where the entire Romani text is translated into Serbian. V: ID_example The unique identifier for each loanword token, combining the token number with a label indicating the source: "AD" for adult speech transcripts or "CH" for elementary-school children's speech transcripts. W: Age group Contains the information on the age group: 1 = adults, 2 = children. X: Source Contains the information on the source of the loanverb token. For the published sources the page number in the source is provided, whereas for the unpublished transcripts the number of the participant is added. The following abbreviations were used for the sources: Rečnik2017 = Ćirković&Mirić 2017, Sikimić2018 = Sikimić 2018, CH = unpublished transcripts of children’s speech. References: Arsenijević, Boban, Gomboc Čeh, Katarin, Marušič, Franc Lanko, Milosavljević, Stefan, Mišmaš, Petra, Simić, Jelena, Simonović, Marko and Žaucer, Rok. 2024. Database of the Western South Slavic Verb HyperVerb 2.0 -- WeSoSlav, Slovenian language resource repository CLARIN.SI, ISSN 2820-4042, http://hdl.handle.net/11356/1846. Ćirković,Svetlana, Mirić, Mirjana. 2017. Romsko-srpski rečnik knjaževačkog gurbetskog govora [Romani-Serbian Dictionary of Knjaževac Gurbet Speech]. Knjaževac: Narodna biblioteka "Njegoš". Mirić, Mirjana, Ćirković, Svetlana. 2022. Gurbetski romski u kontaktu. Analiza balkanizama i pozajmljenica iz srpskog jezika [Gurbet Romani in contact. The Analysis of Balkanisms and Serbian Loanwords]. Belgrade: Institute for Balkan Studies. Sikimić, Biljana (ed.). 2018. Jezik i tradicija knjaževačkih Roma [Language and Tradition of Knjaževac Roma]. Knjaževac: Narodna biblioteka "Njegoš". Acknowledgments: This database is a result of the project “What’s in a verb? Mapping Serbian verbs borrowed into Romani” (Institute for Balkan Studies, Serbian Academy of Sciences and Arts (SASA), Department of Slavic Studies, University of Graz). The project is financed by the Ministry of Science, Technological Development and Innovations of the Republic of Serbia in cooperation with the Austria’s Agency for Education and Internationalisation (OeAD), within the program of scientific and technological cooperation between the Republic of Serbia and the Republic of Austria, for the period 2024–2026.



