遇见数据集

Verbal derivational suffixes in Hungarian: -(s)Odik and -(s)Ul

收藏
Zenodo2023-02-05 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

This is an open-source dataset containing more than 1.1 million corpus occurrences of the Hungarian verbal derivational suffixes -(s)Odik and -(s)Ul. Both suffixes are used to create intransitive verbs from nominal bases, and they both mean 'to become [adjective/noun]'. This dataset is suited for quantitative investigations into the subtle differences regarding how and when these suffixes are used. It consists of the following columns: 1 <em>id</em>: ID 2 <em>form</em>: lowercase word form 3 <em>lemma</em>: word without inflectional suffixes; if the verb has a (separated) preverb, there is a + sign between the preverb and the verb stem 4 <em>prev</em>: preverb associated with the verb 5 <em>prevtype</em>: PFX if the preverb is prefixed to the verb, SEP if the preverb is separated 6 <em>verb</em>: verb lemma; in each case without preverb 7 <em>root</em>: adjective or noun serving as the base of verb formation 8 <em>suffix</em>: derivational suffix: <em>-ul/ül/sul/sül</em> endings are represented by -(s)Ul, <em>-odik/edik/ödik/sodik/sedik/södik</em> endings are represented by -(s)Odik 9 <em>w2v_cluster</em>: the cluster ID of the root, based on word2vec embedding 10 <em>argframe_cases</em>: arguments of the verb, represented by case-endings 11 <em>argframe_long</em>: arguments of the verb, represented by lemma + case-ending combinations 12 <em>doc_year</em>: the year of writing or the year of publication, 0 if unknown 13 <em>doc_style</em>: document style 14 <em>doc_id</em>: document identifier 15 <em>left_context</em>: text preceding the hit 16 <em>kwic</em>: the hit 17 <em>right_context</em>: text following the hit 18 <em>freqsum</em>: token frequency of the verb lemma; occurrences with and without preverbs are counted together 19 <em>prev_vs_all</em>: token frequency of the verb lemma with any preverb, divided by the 'freqsum' value 20 <em>actprev_vs_allprev</em>: token frequency of the specific preverb + verb lemma combination, divided by the 'prev_vs_all' value The first row stands for the header. If a cell's value is unspecified, it is marked with underscore (_).

提供机构:
Zenodo
创建时间:
2023-02-05
二维码
社区交流群
二维码
科研交流群
商业服务