遇见数据集

Annotated dataset of Serbian -nje nominalisations (srWaC)

收藏
Zenodo2026-05-21 更新2026-05-26 收录
官方服务:

资源简介:

Annotated dataset of Serbian -nje nominalisations (srWaC) Marko Simonović (University of Graz) Predrag Kovačević (University of Novi Sad) Goal and rationale This dataset was created as a modern BCMS comparison dataset for the analysis of -nie nominalisations in Simonović, Kovačević and Milićev (accepted). Its purpose is to provide comparable modern BCMS evidence against which the Slavonic-Serbian -nie nominalisation data can be contrasted. Dataset description Each row represents one attested token of a modern BCMS -nje nominalisation drawn from srWaC (Ljubešić and Klubička 2014), a Serbian web corpus collected from the .rs top-level domain and automatically annotated with lemma and other linguistic information. The data were extracted from srWaC on the basis of lemma endings and then manually checked. The sample contains: 1,094 tokens 334 lemmas The data were extracted from the whole srWaC corpus using the following CQL lemma query: [lemma=".*(a|e)nje"] This query targets lemmas ending in -anje or -enje. Since the query targets formal lemma endings, it can retrieve false positives. In this dataset, there are three such false positives. They were not removed, but their status is explicitly marked by 0 in Departicipial and blank values in the subsequent annotation columns. In what follows, we describe each annotated column. Column A — ID ID assigns an arbitrary unique number to each example in the dataset. Column B — Citation form Contains the citation form of the nominalisation, for example proglašenje, udruženje, poboljšanje, interesovanje. Due to morphological richness of the target language, the citation form may differ from the attested surface form in the target column. For example, the citation form proglašenje corresponds to the attested form proglašenja, and udruženje corresponds to Udruženja. Column C — Departicipial Marks whether the nominalisation can be analysed as derived from a passive n-participle. Values: 1 = the item is analysed as a departicipial -nje nominalisation. 0 = the item is not analysed as departicipial. The value 0 is used for false positives retrieved by the extraction query. These items were retained in the dataset but were not further annotated in the columns Compound stem and Perfective base. Examples of false positives marked 0 are kamenje ‘stones’ and Oranje ‘Oranje’. Column D — Compound stem Compound stem marks whether the nominalisation contains a compound stem. Values: 1 = the nominalisation contains a compound stem. 0 = the nominalisation does not contain a compound stem. blank = not applicable because the item was not analysed as a target departicipial nominalisation. For example, e-izdanje is marked as having a compound stem, while forms such as proglašenje, udruženje and interesovanje are marked as non-compound. Note that Slavic prefixes and the negation particle were not counted as compound material. Column E — Perfective base Perfective base marks whether the nominalisation is based on a perfective verb. Values: 1 = the nominalisation has a perfective base. 0 = the nominalisation does not have a perfective base, or the base is imperfective, biaspectual, or otherwise not clearly perfective. blank = not applicable because the item was not analysed as a target departicipial nominalisation. Examples marked 1 include proglašenje, udruženje, poboljšanje, odeljenje, pogubljenje, and objašnjenje. Examples marked 0 include interesovanje, imanje, značenje, pretraživanje, and službovanje. Column F — source Contains the title or identifier of the srWaC text from which the example was extracted. The values appear as article or document titles, for example Odnosi Srbije i Angole, Ušivac, Loner B. I, GNU alati za pretraživanje, and Crkva Svetih apostola Petra i Pavla u Vlasenici. Column G — left context Contains the text immediately preceding the attested target form. It functions as the left side of a concordance line. The context may contain paragraph markers such as ¶, and it may include both Latin and Cyrillic script, reflecting the source material. Column H — target Contains the exact attested surface form of the nominalisation in the corpus. This form may be inflected, capitalised, or otherwise different from the citation form. Column I — right context Contains the text immediately following the attested target form. Together with left context and target, this column provides the concordance context for each example. Reference Ljubešić, Nikola, and Filip Klubička. 2014. “{bs,hr,sr}WaC — Web Corpora of Bosnian, Croatian and Serbian.” In Proceedings of the 9th Web as Corpus Workshop (WaC-9), 29–35. Gothenburg: Association for Computational Linguistics.

提供机构:
Zenodo
创建时间:
2026-05-21
二维码
社区交流群
二维码
科研交流群
商业服务