Cases of Complements of Finnish Verbs
收藏资源简介:
Context Cases of the complements of Finnish verbs. The data is useful for natural language generation (NLG). The data is described in the following paper, which should also be cited if this data is used: Hämäläinen, Mika and Rueter, Jack 2018. Development of an Open Source Natural Language Generation Tool for Finnish. In Proceedings of the Fourth International Workshop on Computational Linguistics of Uralic Languages, 51–58. Content The file contains a list of Finnish verbs from the Finnish Internet Parsebank, these have been lemmatized and filtered by part-of-speech with Omorfi. For each verb, there is a list of grammatical cases together with how many times that case has occurred with a syntactic relation to the verb on the right context of the verb. Essentially, this data can be used to see the most typical direct object case (partitive, genitive, elative..) for each Finnish verb. The data can also indicate whether the verb can take an indirect object as well or not. Inspiration This data is important for NLG tasks. One could learn to predict if a verb can take a direct object or also an indirect object. This data has been used to generate poems in Finnish: Hämäläinen, M. (2018). Harnessing NLG to Create Finnish Poetry Automatically. In F. Pachet, A. Jordanous, & C. León (Eds.), Proceedings of the Ninth International Conference on Computational Creativity (pp. 9-15). Salamanca: Association for Computational Creativity (ACC).
本数据集聚焦芬兰语动词补语的语境格位。 本数据集可用于自然语言生成(Natural Language Generation,以下简称NLG)任务。若使用本数据集,请一并引用以下文献:Hämäläinen, Mika 与 Rueter, Jack(2018). 《面向芬兰语的开源自然语言生成工具开发》(Development of an Open Source Natural Language Generation Tool for Finnish),载于《第四届乌拉尔语系计算语言学国际研讨会论文集》(Proceedings of the Fourth International Workshop on Computational Linguistics of Uralic Languages),第51-58页。 ## 数据集内容 本文件收录了源自芬兰语互联网语料库(Finnish Internet Parsebank)的芬兰语动词列表,所有动词均通过Omorfi工具完成词形还原与词性过滤。针对每个动词,本数据集均附带一份语法格位列表,并标注了该格位与动词在右侧语境中形成句法关系的出现频次。本质而言,本数据集可用于探究每个芬兰语动词最典型的直接宾语格位(如部分格、属格、离格等),同时可判断该动词是否可接间接宾语。 ## 应用价值 本数据集对NLG任务具有重要意义,研究者可基于本数据集训练模型,预测动词是否可接直接宾语或间接宾语。本数据集已被用于自动生成芬兰语诗歌:Hämäläinen, M.(2018). 《利用自然语言生成技术自动创作芬兰语诗歌》(Harnessing NLG to Create Finnish Poetry Automatically),载于F. Pachet、A. Jordanous与C. León编辑的《第九届计算创造力国际会议论文集》(Proceedings of the Ninth International Conference on Computational Creativity),第9-15页,萨拉曼卡:计算创造力协会(Association for Computational Creativity,ACC)。



