遇见数据集

Annotated dataset of Russian -nie nominalisations

收藏
Zenodo2026-05-22 更新2026-05-26 收录
官方服务:

资源简介:

Annotated dataset of Russian -nie nominalisations Daria Seres (University of Graz) Marko Simonović (University of Graz) Predrag Kovačević (University of Novi Sad) Goal and rationale This dataset was created as a modern Russian comparison dataset for the analysis of -nie nominalisations in Simonović, Kovačević and Milićev (accepted). Its purpose is to provide comparable modern Russian evidence against which the Slavonic-Serbian -nie nominalisation data can be contrasted. Dataset description Each row represents one attested token of a modern Russian -nie nominalisation drawn from the Russian National Corpus (RNC; ruscorpora.ru). The data were extracted from RNC texts classified as учебно-научная ‘academic/educational-scientific’ and публицистика ‘journalistic writing’ and restricted to texts produced after 2000. The sample was randomly selected and then manually checked and annotated. The sample contains: 1,705 tokens 435 lemmas The data were extracted from the Russian National Corpus using the following lemma query: [lemma=".*(т|н)ие"] This query targets lemmas ending in -ние or -тие. Since the query targets formal lemma endings, it can retrieve false positives. In this dataset, such items were not removed, but their status is explicitly marked in Departicipial and by blank or uncertain values in the subsequent annotation columns. In what follows, we describe each column. The columns Departicipial, Compound stem, and Perfective base contain manual linguistic annotation; the remaining columns contain identifiers, concordance context, or metadata from the Russian National Corpus. In what follows, we describe each column. The columns Departicipial, Compound stem, and Perfective base contain manual linguistic annotation; the remaining columns contain identifiers, concordance context, or metadata from the Russian National Corpus. Column A — ID ID assigns an arbitrary unique number to each example in the dataset. Column B — Citation form Contains the citation form of the nominalisation, for example разрешение ‘permission’, клонирование ‘cloning’, исследование ‘research’, образование ‘education’, исполнение ‘execution; fulfilment’, and создание ‘creation’. Due to Russian inflectional morphology, the citation form may differ from the attested surface form in the Center column. For example, the citation form клонирование ‘cloning’ corresponds to attested forms such as клонирования, and разрешение ‘permission’ may correspond to разрешение or разрешения depending on case and number. Column C — Departicipial Marks whether the item is analysed as belonging to the target deverbal/departicipial -nie nominalisation class. Values: 1 = the item is analysed as a target deverbal/departicipial -nie nominalisation. 0 = the item is not analysed as a target deverbal/departicipial -nie nominalisation. ? = the status is uncertain. Examples marked 1 include разрешение ‘permission’, клонирование ‘cloning’, использование ‘use’, создание ‘creation’, исследование ‘research’, and выполнение ‘performance; fulfilment’. Items that do not have an attested corresponding verb, but for which that verb could be constructed and its aspect could be determined were also marked 1. Examples marked 0 include formal false positives such as тысячелетие ‘millennium’, 200-летие ‘200th anniversary’, 850-летие ‘850th anniversary’, поколение ‘generation’ and Минобразования ‘Ministry of Education’. These items match the formal extraction pattern but were not analysed as target deverbal/departicipial nominalisations. Examples marked ? include uncertain cases where the nominal is perceived as deverbal/departicipial, but the corresponding verb could not be constructed. Examples are волнение ‘agitation’ and возникновение ‘appearance, emergence’. These were retained in the dataset but marked as uncertain rather than forced into a binary classification. Column D — Compound stem Compound stem marks whether the nominalisation contains a compound stem. Values: 1 = the nominalisation contains a compound stem. 0 = the nominalisation does not contain a compound stem. blank = not applicable or not annotated, usually because the item was not analysed as a target nominalisation or its status was uncertain. Examples marked 1 include налогообложение ‘taxation’, правоотношение ‘legal relation’, здравоохранение ‘health care’, водоснабжение ‘water supply’, телевидение ‘television’, and мироощущение ‘worldview; sense of the world’. Examples marked 0 include разрешение ‘permission’, клонирование ‘cloning’, исследование ‘research’, использование ‘use’, создание ‘creation’, and решение ‘decision; solution’. Note that Slavic prefixes and the negation particle were not counted as compound material. Column E — Perfective base Perfective base marks whether the nominalisation is based on a perfective verb. Values: 1 = the nominalisation has a perfective base. 0 = the nominalisation does not have a perfective base, or the base is imperfective, biaspectual, lexicalised, or otherwise not clearly perfective. ? = the aspectual status of the base is uncertain. blank = not applicable or not annotated, usually because the item was not analysed as a target nominalisation or its status was uncertain. Examples marked 1 include разрешение ‘permission’, воссоздание ‘recreation’, создание ‘creation’, окончание ‘ending; completion’, возрождение ‘revival’, выполнение ‘performance; fulfilment’, сокращение ‘reduction’, and повышение ‘increase’. Examples marked 0 include клонирование ‘cloning’, исследование ‘research’, использование ‘use’, образование ‘education’, страхование ‘insurance’, течение ‘course; flow’, содержание ‘content; maintenance’, and соревнование ‘competition’. Column F — Unattested verb Marks whether the base verb is an actually attested verb. For instance, разрешение ‘permission’ has the value 1 because разрешить is an attested verb, but упражнение ‘exercise’ got the value 0 because упражнить is not an attested verb, even though it can be constructed and its aspect can be determined. Column G — Left context Contains the text immediately preceding the attested target form. It functions as the left side of a concordance line. For example, in the row with разрешение ‘permission’, the left context is: мамонта, японской стороне предстоит получить специальное Column H — Center Contains the exact attested surface form of the nominalisation in the corpus. This form may be inflected, capitalised, or otherwise different from the citation form. For example, the citation form растение ‘plant’ corresponds to attested forms such as растений, and решение ‘decision; solution’ corresponds to forms such as решение, решения, and решений. Column I — Right context Contains the text immediately following the attested target form. Together with Left context and Center, this column provides the concordance context for each example. For example, in the row with разрешение ‘permission’, the right context is: от российских властей, что предположительно будет Column J — Title Contains the title or bibliographic identifier of the RNC source text from which the example was extracted. Examples include Японцы хотят клонировать якутского мамонта ‘The Japanese want to clone a Yakut mammoth’, Юрий Патрикеев: «Проиграл, потому что боролся с Гарднером слишком долго», and Учет и налогообложение операций по страхованию работников ‘Accounting and taxation of employee insurance operations’. Column K — Author Contains the author name recorded in the Russian National Corpus metadata, where available. Column L — Birthday Contains the author’s year of birth where this information is available in the corpus metadata. The field is blank when the information is not provided. Column M — Header Contains the document header recorded in the RNC metadata. In many cases this repeats or shortens the source title. Column N — Created Contains the year in which the source text was created. The dataset is restricted to texts produced after 2000. Column O — Sphere Contains the broad RNC domain classification. In this dataset, the relevant values are учебно-научная ‘academic/educational-scientific’ and публицистика ‘journalistic writing’. Column P — Type Contains the text-type classification assigned by the RNC. Examples in the dataset include заметка ‘short article/note’, интервью ‘interview’, статья ‘article’, аннотация ‘abstract’, письмо деловое ‘business letter’, and характеристика ‘character reference/evaluation’. Column Q — Topic Contains the thematic classification assigned by the RNC. Examples include наука и технологии ‘science and technology’, искусство и культура ‘art and culture’, политика и общественная жизнь ‘politics and public life’, право ‘law’, образование ‘education’, спорт ‘sport’, and бизнес, коммерция, экономика, финансы ‘business, commerce, economics, finance’. Column R — Publication Contains the publication source recorded in the RNC metadata. Examples include «Известия», «Домовой», «Вечерняя Москва», «Вопросы статистики», «Физика твердого тела», and «Бухгалтерский учёт». Column S — Publ_year Contains the publication year recorded in the RNC metadata. This usually corresponds to the year in Created, but the two fields are kept separate because they come from distinct metadata fields. Column T — Medium Contains the medium classification assigned by the RNC. Examples include газета ‘newspaper’, журнал ‘magazine/journal’, машинопись ‘typescript’, and электронный текст ‘electronic text’. Column U — Ambiguity Contains information on ambiguity resolution in the RNC metadata. In the present dataset, examples commonly have the value омонимия снята ‘homonymy resolved’, indicating that morphological ambiguity was resolved in the corpus annotation. Column V — Full context Contains the complete corpus context from which the example was extracted. This field allows the reader to verify the local interpretation of the nominalisation and its annotation. For example, the full context for разрешение ‘permission’ is: Для того чтобы вывезти в Японию фрагмент ткани мамонта, японской стороне предстоит получить специальное разрешение от российских властей, что предположительно будет сделано осенью. Reference Russian National Corpus. Available at: ruscorpora.ru. Simonović, Marko, Predrag Kovačević and Tanja Milićev. Accepted. When -nie met -nje: Slavonic-Serbian loan deverbal nominals. To appear in Stefan Milosavljević, Daria Seres, Jelena Stojković, Marko Simonović and Jelena Živojinović (eds.), Advances in Formal Slavic Linguistics 2023. Berlin: Language Science Press.

俄语-ние派生词标注数据集 作者: Daria Seres(格拉茨大学) Marko Simonović(格拉茨大学) Predrag Kovačević(诺威萨德大学) 研究目标与理论依据 本数据集为适配Simonović、Kovačević与Milićev(已录用)中关于俄语-ние派生词的分析需求而构建,属于现代俄语对比数据集。其核心目的是提供可对比的现代俄语语料,用于与斯拉夫语-塞尔维亚语的-ние派生词数据进行对照分析。 数据集说明 本数据集的每一行对应一条取自俄罗斯国家语料库(Russian National Corpus,简称RNC;网址ruscorpora.ru)的已证实的现代俄语-ние派生词实例。语料取自RNC中归类为"学术/教育-科学"(учебно-научная)与"新闻写作"(публицистика)的文本,且限定为2000年之后创作的文本。样本经随机抽样后,由人工进行核查与标注。 本数据集共包含: 1705个词例 435个词元(lemma) 本数据集通过以下词元查询语句从RNC中提取语料:`[lemma=".*(т|н)ие"]`。该查询语句针对以-ние或-тие结尾的词元。由于仅依据形式上的词元结尾进行匹配,可能会提取出误判项。本数据集未移除此类误判项,而是在"分词派生属性(Departicipial)"列中明确标注其状态,并在后续标注列中留空或使用不确定值。 下文将逐一介绍各列。其中,分词派生属性(Departicipial)、复合词干属性(Compound stem)与完成体词基属性(Perfective base)三列为人工语言标注;其余列则存储标识符、索引上下文或来自俄罗斯国家语料库的元数据。 下文将逐一介绍各列。其中,分词派生属性(Departicipial)、复合词干属性(Compound stem)与完成体词基属性(Perfective base)三列为人工语言标注;其余列则存储标识符、索引上下文或来自俄罗斯国家语料库的元数据。 A列 — 编号(ID):为数据集中的每个实例分配唯一的任意编号。 B列 — 词元形式(Citation form):存储该派生词的词元形式,例如:разрешение(许可)、клонирование(克隆)、исследование(研究)、образование(教育)、исполнение(执行;履行)以及создание(创建)。由于俄语屈折形态的影响,词元形式可能与H列中的实采表面形式存在差异。例如,词元形式клонирование(克隆)对应的实采形式可为клонирования,而разрешение(许可)则可根据格与数的变化对应разрешение或разрешения。 C列 — 分词派生属性(Departicipial):标注该词项是否被归为目标动转名/分词派生-ние派生词类别。取值说明: 1 = 该词项被分析为目标动转名/分词派生-ние派生词; 0 = 该词项未被归为目标动转名/分词派生-ние派生词; ? = 状态不确定。 被标记为1的示例包括разрешение(许可)、клонирование(克隆)、использование(使用)、создание(创建)、исследование(研究)以及выполнение(执行;履行)。即便某类派生词暂无已证实的对应动词,但可构建其对应动词且可确定体貌的,同样标记为1。 被标记为0的示例包括形式上符合提取规则的误判项,如тысячелетие(千年)、200-летие(200周年纪念)、850-летие(850周年纪念)、поколение(世代)以及Минобразования(教育部)——此类词项虽符合形式提取模式,但未被归为目标动转名/分词派生派生词。 被标记为?的示例为存疑情况:该派生词可被视为动转名/分词派生,但无法构建其对应动词,例如волнение(骚动)与возникновение(出现;产生)。此类词项保留在数据集中,但标记为不确定状态,而非强制归入二分类。 D列 — 复合词干属性(Compound stem):标注该派生词是否包含复合词干。取值说明: 1 = 该派生词包含复合词干; 0 = 该派生词不包含复合词干; 空白 = 不适用或未标注,通常因该词项未被归为目标派生词或状态不确定。 标记为1的示例包括налогообложение(征税)、правоотношение(法律关系)、здравоохранение(医疗保健)、водоснабжение(供水)、телевидение(电视)以及мироощущение(世界观;对世界的感知)。 标记为0的示例包括разрешение(许可)、клонирование(克隆)、исследование(研究)、использование(使用)、создание(创建)以及решение(决定;解决方案)。 注意:斯拉夫语前缀与否定小品词不计入复合词干成分。 E列 — 完成体词基属性(Perfective base):标注该派生词的词基是否为完成体动词。取值说明: 1 = 该派生词的词基为完成体动词; 0 = 该派生词的词基为非完成体、双体、词汇化或其他无法明确为完成体的情况; ? = 词基的体貌状态不确定; 空白 = 不适用或未标注,通常因该词项未被归为目标派生词或状态不确定。 标记为1的示例包括разрешение(许可)、воссоздание(再造)、создание(创建)、окончание(结尾;完成)、возрождение(复兴)、выполнение(执行;履行)、сокращение(缩减)以及повышение(提升)。 标记为0的示例包括клонирование(克隆)、исследование(研究)、использование(使用)、образование(教育)、страхование(保险)、течение(进程;流动)、содержание(内容;维护)以及соревнование(竞赛)。 F列 — 未证实动词属性(Unattested verb):标注该派生词的词基动词是否为已证实的动词。例如,разрешение(许可)的取值为1,因为разрешить是已证实的动词;而упражнение(练习)的取值为0,因为упражнить并非已证实的动词,即便可构建该动词且可确定其体貌。 G列 — 左侧上下文(Left context):存储实采目标形式紧邻的前文文本,作为索引行的左侧部分。例如,包含разрешение(许可)的行的左侧上下文为:`мамонта, японской стороне предстоит получить специальное` H列 — 目标形式(Center):存储语料库中实采的该派生词的精确表面形式。该形式可能存在屈折变化、大写形式或其他与词元形式不同的情况。例如,词元形式растение(植物)对应的实采形式可为растений,而решение(决定;解决方案)对应的实采形式可为решение、решения或решений。 I列 — 右侧上下文(Right context):存储实采目标形式紧邻的后文文本。与左侧上下文、目标形式共同构成每个实例的索引上下文。例如,包含разрешение(许可)的行的右侧上下文为:`от российских властей, что предположительно будет` J列 — 标题(Title):存储提取该实例的RNC源文本的标题或文献标识符。示例包括:《Японцы хотят клонировать якутского мамонта》(日本人希望克隆雅库特猛犸象)、《Юрий Патрикеев: «Проиграл, потому что боролся с Гарднером слишком долго»》(尤里·帕特里凯耶夫:"我输了,因为和加德纳缠斗太久")以及《Учет и налогообложение операций по страхованию работников》(员工保险业务的核算与征税)。 K列 — 作者(Author):存储RNC元数据中记录的作者姓名(如可获取)。 L列 — 出生年份(Birthday):存储语料库元数据中可获取的作者出生年份,若未提供相关信息则字段留空。 M列 — 文档页眉(Header):存储RNC元数据中记录的文档页眉,多数情况下会重复或简化源文本标题。 N列 — 创建年份(Created):存储源文本的创作年份。本数据集限定使用2000年之后创作的文本。 O列 — 领域分类(Sphere):存储RNC的宽泛领域分类。本数据集中的有效取值为"学术/教育-科学"(учебно-научная)与"新闻写作"(публицистика)。 P列 — 文本类型(Type):存储RNC分配的文本类型分类。示例包括:заметка(短文/笔记)、интервью(访谈)、статья(文章)、аннотация(摘要)、письмо деловое(商务信函)以及характеристика(推荐信/评价)。 Q列 — 主题分类(Topic):存储RNC分配的主题分类。示例包括:наука и технологии(科学与技术)、искусство и культура(艺术与文化)、политика и общественная жизнь(政治与公共生活)、право(法律)、образование(教育)、спорт(体育)以及бизнес, коммерция, экономика, финансы(商业、贸易、经济与金融)。 R列 — 出版来源(Publication):存储RNC元数据中记录的出版来源。示例包括:《Известия》(《消息报》)、《Домовой》(《家》)、《Вечерняя Москва》(《莫斯科晚报》)、《Вопросы статистики》(《统计学问题》)、《Физика твердого тела》(《固体物理》)以及《Бухгалтерский учёт》(《会计》)。 S列 — 出版年份(Publ_year):存储RNC元数据中记录的出版年份。该年份通常与创建年份(Created)一致,但因二者来自不同的元数据字段,故保留为独立列。 T列 — 传播媒介(Medium):存储RNC分配的媒介分类。示例包括:газета(报纸)、журнал(期刊/杂志)、машинопись(打印稿)以及электронный текст(电子文本)。 U列 — 歧义处理(Ambiguity):存储RNC元数据中关于歧义消解的信息。本数据集中的示例通常取值为"омимия снята"(歧义已消解),表示语料库标注中已解决了形态歧义问题。 V列 — 完整上下文(Full context):存储提取该实例的完整语料库上下文,便于读者验证该派生词的本地释义及其标注情况。例如,包含разрешение(许可)的完整上下文为:`Для того чтобы вывезти в Японию фрагмент ткани мамонта, японской стороне предстоит получить специальное разрешение от российских властей, что предположительно будет сделано осенью.`(为将猛犸象织物碎片运往日本,日方需获得俄罗斯官方的特殊许可,预计该许可将于秋季获批。) 参考文献 俄罗斯国家语料库. 可访问地址:https://ruscorpora.ru. Simonović, Marko, Predrag Kovačević 与 Tanja Milićev. 已录用. 当-ние遇上-ње:斯拉夫语-塞尔维亚语借入动转名派生词. 收录于Stefan Milosavljević、Daria Seres、Jelena Stojković、Marko Simonović与Jelena Živojinović主编,《2023年形式斯拉夫语言学进展》(Advances in Formal Slavic Linguistics 2023). 柏林:语言科学出版社(Language Science Press),待刊。

提供机构:
Zenodo
创建时间:
2026-05-22
二维码
社区交流群
二维码
科研交流群
商业服务