遇见数据集

Annotated dataset of Serbian -nje nominalisations (CLASSLAWIKI-sr)

收藏
Zenodo2026-05-21 更新2026-05-26 收录
官方服务:

资源简介:

Annotated dataset of Serbian -nje nominalisations (CLASSLAWIKI-sr) Marko Simonović (University of Graz) Predrag Kovačević (University of Novi Sad) Goal and rationale This dataset was created as a modern BCMS comparison dataset for the analysis of -nie nominalisations in Simonović, Kovačević and Milićev (accepted). Its purpose is to provide comparable modern BCMS evidence against which the Slavonic-Serbian -nie nominalisation data can be contrasted. Dataset description Each row represents one attested token of a modern BCMS -nje nominalisation drawn from CLASSLAWiki-sr (Ljubešić et al. 2021), the Serbian component of the CLASSLA-Wikipedia 1.0 collection, based on Serbian Wikipedia and automatically annotated with lemmas and other linguistic information. The data were extracted from CLASSLAWIKI-sr on the basis of lemma endings and then manually checked. The sample contains: 1,094 tokens 334 lemmas The data were extracted from the whole CLASSLAWIKI-sr corpus using the following CQL lemma query: [lemma=".*(a|e)nje"] This query targets lemmas ending in -anje or -enje. Since the query targets formal lemma endings, it can retrieve false positives. In this dataset, there are three such false positives. They were not removed, but their status is explicitly marked by 0 in Departicipial and blank values in the subsequent annotation columns. In what follows, we describe each annotated column. Column A — ID ID assigns an arbitrary unique number to each example in the dataset. Column B — Citation form Contains the citation form of the nominalisation, for example proglašenje, udruženje, poboljšanje, interesovanje. Due to morphological richness of the target language, the citation form may differ from the attested surface form in the target column. For example, the citation form proglašenje corresponds to the attested form proglašenja, and udruženje corresponds to Udruženja. Column C — Departicipial Marks whether the nominalisation can be analysed as derived from a passive n-participle. Values: 1 = the item is analysed as a departicipial -nje nominalisation. 0 = the item is not analysed as departicipial. The value 0 is used for false positives retrieved by the extraction query. These items were retained in the dataset but were not further annotated in the columns Compound stem and Perfective base. Examples of false positives marked 0 are kamenje ‘stones’ and Oranje ‘Oranje’. Column D — Compound stem Compound stem marks whether the nominalisation contains a compound stem. Values: 1 = the nominalisation contains a compound stem. 0 = the nominalisation does not contain a compound stem. blank = not applicable because the item was not analysed as a target departicipial nominalisation. For example, e-izdanje is marked as having a compound stem, while forms such as proglašenje, udruženje and interesovanje are marked as non-compound. Note that Slavic prefixes and the negation particle were not counted as compound material. Column E — Perfective base Perfective base marks whether the nominalisation is based on a perfective verb. Values: 1 = the nominalisation has a perfective base. 0 = the nominalisation does not have a perfective base, or the base is imperfective, biaspectual, or otherwise not clearly perfective. blank = not applicable because the item was not analysed as a target departicipial nominalisation. Examples marked 1 include proglašenje, udruženje, poboljšanje, odeljenje, pogubljenje, and objašnjenje. Examples marked 0 include interesovanje, imanje, značenje, pretraživanje, and službovanje. Column F — source Contains the title or identifier of the CLASSLAWIKI-sr text from which the example was extracted. The values appear as article or document titles, for example Odnosi Srbije i Angole, Ušivac, Loner B. I, GNU alati za pretraživanje, and Crkva Svetih apostola Petra i Pavla u Vlasenici. Column G — left context Contains the text immediately preceding the attested target form. It functions as the left side of a concordance line. The context may contain paragraph markers such as ¶, and it may include both Latin and Cyrillic script, reflecting the source material. Column H — target Contains the exact attested surface form of the nominalisation in the corpus. This form may be inflected, capitalised, or otherwise different from the citation form. Column I — right context Contains the text immediately following the attested target form. Together with left context and target, this column provides the concordance context for each example. Reference Ljubešić, Nikola, Filip Markoski, Elena Markoska & Tomaž Erjavec. 2021. Comparable corpora of South-Slavic Wikipedias CLASSLA-Wikipedia 1.0. Jožef Stefan Institute. http://hdl.handle.net/11356/1427

塞尔维亚语-nje型名词化标注数据集(CLASSLAWIKI-sr) 马尔科·西莫诺维奇(格拉茨大学) 普雷德拉格·科瓦切维奇(诺维萨德大学) 研究目标与理论依据 本数据集为现代巴尔干克罗地亚语-塞尔维亚语-黑山语(BCMS)对比数据集,用于分析西蒙诺维奇、科瓦切维奇与米利切维奇(已录用)研究中的-nje型名词化现象。其旨在提供可对比的现代BCMS语料证据,用于与斯拉夫-塞尔维亚语-nje型名词化数据进行对照分析。 数据集说明 每一行对应一条从CLASSLAWiki-sr(Ljubešić等,2021)中提取的现代BCMS -nje型名词化实见语料单元(token)。CLASSLAWiki-sr是CLASSLA-Wikipedia 1.0语料库的塞尔维亚语子库,基于塞尔维亚语维基百科构建,并已自动标注词元(lemma)与其他语言信息。本数据集基于词元后缀从CLASSLAWIKI-sr中提取,随后经人工核查。 本样本集包含: 1094个语料单元 334个词元 研究人员通过如下语料库查询语言(CQL)查询语句从完整的CLASSLAWIKI-sr语料库中提取数据: [lemma=".*(a|e)nje"] 该查询旨在提取以-anje或-enje结尾的词元,但由于仅基于形式化词元后缀进行匹配,可能会引入误判条目。本数据集中共存在3条此类误判样本,未将其移除,而是通过"分词派生(Departicipial)"列标记为0,并在后续标注列留空以明确其状态。 下文将逐一说明各标注列的含义: 列A — ID ID列:为数据集中的每个样本分配唯一任意编号。 列B — 引述形式 引述形式列:存储该名词化形式的标准引述形式,例如"proklašenje"、"udruženje"、"poboljšanje"、"interesovanje"。由于目标语言形态丰富,引述形式可能与目标列中的实见表层形式存在差异。例如,引述形式"proklašenje"对应实见形式"proklašenja","udruženje"对应"Udruženja"。 列C — "分词派生(Departicipial)" 分词派生列:标记该名词化是否可分析为源自被动过去分词的派生形式。 取值说明: 1 = 该条目可归类为分词派生型-nje名词化 0 = 该条目不属于此类 其中0值用于标记提取查询返回的误判样本,此类条目虽保留在数据集中,但未在复合词干与完成体词根列中进行进一步标注。标记为0的误判样本示例包括"kamenje"(意为“石块”)与"Oranje"(意为“奥兰治”)。 列D — 复合词干 复合词干列:标记该名词化是否包含复合词干。 取值说明: 1 = 该名词化包含复合词干 0 = 该名词化不包含复合词干 留空 = 不适用,因该条目未被归类为目标分词派生型名词化 需注意,斯拉夫语前缀与否定小品词不计入复合成分。示例:"e-izdanje"被标记为包含复合词干,而"proklašenje"、"udruženje"与"interesovanje"等形式则被标记为非复合形式。 列E — 完成体词根 完成体词根列:标记该名词化是否基于完成体动词。 取值说明: 1 = 该名词化以完成体动词为词根 0 = 该名词化不以完成体动词为词根(或词根为未完成体、双体动词,或完成体属性不明确) 留空 = 不适用,因该条目未被归类为目标分词派生型名词化 标记为1的示例包括"proklašenje"、"udruženje"、"poboljšanje"、"odeljenje"、"pogubljenje"与"objašnjenje";标记为0的示例包括"interesovanje"、"imanje"、"značenje"、"pretraživanje"与"službovanje"。 列F — 来源 来源列:存储提取该样本的CLASSLAWIKI-sr文本的标题或标识符,取值为文章或文档标题,例如"Odnosi Srbije i Angole"、"Ušivac"、"Loner B. I"、"GNU alati za pretraživanje"与"Crkva Svetih apostola Petra i Pavla u Vlasenici"。 列G — 左上下文 左上下文列:存储实见目标形式紧邻的前文文本,对应关键词在上下文(KWIC)行的左侧部分。上下文可能包含段落标记¶,同时可包含拉丁字母与西里尔字母,以反映源语料的原始格式。 列H — 目标 目标列:存储语料库中实见的该名词化形式的精确表层形式,该形式可能存在屈折变化、大写格式差异,或与引述形式存在其他不同。 列I — 右上下文 右上下文列:存储实见目标形式紧邻的后文文本。与左上下文及目标列共同构成每个样本的KWIC上下文信息。 参考文献 Ljubešić, Nikola, Filip Markoski, Elena Markoska 与 Tomaž Erjavec. 2021. 南斯拉夫语维基百科可比语料库CLASSLA-Wikipedia 1.0. 约泽夫·斯特凡研究所. http://hdl.handle.net/11356/1427

提供机构:
Zenodo
创建时间:
2026-05-21
二维码
社区交流群
二维码
科研交流群
商业服务