遇见数据集

Lec-nominalizations with an adjusted secondary imperfective morpheme in Slovenian

收藏
Zenodo2024-12-04 更新2026-05-26 收录
官方服务:

资源简介:

Lec-nominalizations with an adjusted secondary imperfective morpheme in Slovenian This dataset is a derivative of Arsenijević et al. (2024). The added columns are E, F, G, H, I, and J. All other columns are retained from the original dataset, and the original source should be referenced for them. Data collection and annotation procedure The goal of the data collection is to identify Slovenian lec-nominalizations (in the original dataset listed as lc-) that have an adjustment of the secondary imperfectivizing morpheme not attested in the corresponding verb. For instance, the verb obračunavati ‘to calculate’ contains the secondary imprefectivizer -av-, which gets adjusted to -ov- in the corresponding lec-nominalization: obračunovalec ‘calculator’. As a rule, the adjustment is always in the direction of -ov- (which has the allomorph -ev- after a set of consonants). To obtain all relevant nominalizations, the national corpus Gigafida 2.0 was searched for nominalizations ending in -ovalec and -evalec using the regular expression [lemma=".*(e|o)valec"]. The search results were then cross-referenced with the verbs listed in Arsenijević et al. (2024), and all lec-nominalizations containing an adjusted version of a verb from Arsenijević et al. (2024) were extracted. For each verb with a corresponding nominalization involving an adjusted secondary imperfectivizer, separate corpus searches were conducted. These searches focused on: The lec-nominalization without adjustment (e.g., obračunavalec), regardless of whether that nominalization is marked as possible in Arsenijević et al. (2024). All attested -lec nominalizations with an adjustment of the secondary imperfective suffix (e.g., obračunovalec). Summary of the columns The results are in 3 of the added columns, 3 columns give frequencies of each individual lec-nominalization: E: regular lc Lists the expected lec-nominalization without adjustment, formed by replacing the inflectional ending -ti from the citation form of the verb with -lec. F: frequency regular lc Contains the frequency of the lec-nominalization without adjustment in Gigafida 2.0. The CQL targeted the specific lemma. For instance, the CQL used for obračunavalec was [lemma="obračunavalec"]. G: lc with adjustment 1 Lists the first lec-nominalization with an adjustment derived from the verb in question and attested in Gigafida 2.0. Given the procedure described above, all items listed in this column contain an -ov-/-ev- that is not present in the base verb. H: frequency lc with adjustment 1 Contains the frequency of the lec-nominalization with an adjustment in Gigafida 2.0. The CQL targeted the specific lemma. For instance, the CQL used for obračunovalec was [lemma="obračunovalec"]. I: lc with adjustment 2 Lists the second lec-nominalization with an adjustment derived from the verb in question and attested in Gigafida 2.0 (if available). J: frequency lc with adjustment 2 Contains the frequency of the lec-nominalization with an adjustment in Gigafida 2.0. The CQL targeted the specific lemma. References Arsenijević, B., Marušič, F. L., Milosavljević, S., Mišmaš, P., Simonovic, M., & Žaucer, R. (2024). Database of the Western South Slavic Verb HyperVerb (WeSoSlaV) – Deverbal Nominalizations [Data set]. Zenodo. https://doi.org/10.5281/zenodo.14230589

斯洛文尼亚语中带有调整后次级未完成体化语素的lec派生词(lec-nominalizations) 本数据集为Arsenijević等人(2024)研究成果的衍生数据集,新增E、F、G、H、I、J共6列字段,其余列均保留自原始数据集,相关内容的溯源需参考原始数据集来源。 数据收集与标注流程 本数据集的收集目标为识别斯洛文尼亚语中的lec派生词(lec-nominalizations,原始数据集中标注为lc-),这类派生词的次级未完成体化语素(secondary imperfectivizing morpheme)经过了调整,且该调整未在对应动词中出现。举例而言,动词obračunavati意为“计算”,其本身包含次级未完成体化后缀-av-,而对应的lec派生词中该后缀被调整为-ov-,即obračunovalec(意为“计算器”)。通常而言,这类调整均指向-ov-(在特定辅音后会出现同位异音形式-ev-)。 为获取所有相关派生词,研究人员使用正则表达式[lemma=".*(e|o)valec"]在国家语料库Gigafida 2.0中检索以-ovalec和-evalec结尾的派生词。随后将检索结果与Arsenijević等人(2024)中列出的动词进行交叉比对,提取出所有源自该文献中动词、且带有调整后形式的lec派生词。 针对每一个带有经调整次级未完成体化语素的派生词对应的动词,研究人员开展了独立语料库检索,检索范围包括: 1. 未经调整的lec派生词(例如obračunavalec),无论该派生词是否在Arsenijević等人(2024)中被标记为可行形式; 2. 所有经证实的、带有次级未完成体后缀调整的-lec派生词(例如obračunovalec)。 列字段说明 新增字段的统计结果分布于3个列中,另有3列用于记录各类lec派生词的出现频次: - E:标准lc形式 列出未经调整的标准lec派生词,其构词方式为将动词原形的屈折结尾-ti替换为-lec。 - F:标准lc形式频次 记录该未经调整的lec派生词在Gigafida 2.0语料库中的出现频次。检索所用的语料库查询语言(Corpus Query Language,简称CQL)针对特定词元,例如针对obračunavalec的检索式为[lemma="obračunavalec"]。 - G:经调整1次的lc形式 列出首个源自对应动词、且在Gigafida 2.0中得到证实的带调整的lec派生词。根据前述检索流程,该列中所有条目均包含基础动词中未出现的-ov-/-ev-后缀。 - H:经调整1次的lc形式频次 记录该带调整的lec派生词在Gigafida 2.0语料库中的出现频次。检索所用CQL针对特定词元,例如针对obračunovalec的检索式为[lemma="obračunovalec"]。 - I:经调整2次的lc形式 列出第二个源自对应动词、且在Gigafida 2.0中得到证实的带调整的lec派生词(若存在)。 - J:经调整2次的lc形式频次 记录该带调整的lec派生词在Gigafida 2.0语料库中的出现频次。检索所用CQL针对特定词元。 参考文献 Arsenijević, B., Marušič, F. L., Milosavljević, S., Mišmaš, P., Simonovic, M., & Žaucer, R. (2024). 西南斯拉夫语动词超动词数据库(WeSoSlaV)——派生词转名词化[数据集]. Zenodo. https://doi.org/10.5281/zenodo.14230589

提供机构:
Zenodo
创建时间:
2024-12-04
二维码
社区交流群
二维码
科研交流群
商业服务