遇见数据集

Norsk Ordbank - Norwegian Nynorsk 2012

收藏
data.europa2022-02-01 更新2025-06-01 收录
官方服务:

资源简介:

Norsk Ordbank - Nynorsk is a lexical database reflecting the official spelling reform that took effect on 1 August 2012, and later adjustments to the spelling for Norwegian Nynorsk. The database consists of a basic vocabulary (lemmas) and a set of inflectional patterns. Each lemma has one or more inflectional patterns connected with it. Each inflectional pattern contains a number of lines that spells out every inflected form of the lemma. Each line contains a transformation pattern and information about word class and morphological features. The pattern shows how the base word can be expanded into an inflected form. The data is stored in seven tables. The table "lemma" contains all entries in Nynorskordboka (an official Nynorsk dictionary) with specification of article number. The list of full forms contains all possible inflected forms of all entries, in accord with current official spelling. (Note that this table also contains forms that are a result of overgeneration, e.g. the putative plural form 'snøar' ('snows') of 'snø' ('snow'). The tables "lemma_paradigme", "paradigme", "paradigme_boying", "boyingsgruppe" and "boying" contain information that is necessary to generate the full forms based on the basic vocabulary ("lemma"). In other words, they contain the link between the lemmas and inflectional patterns, rules and categorial information. The table "leddanalyse" contains information on the decomposition of compund words. In Nynorskordboka, decomposition is indicated with a vertical line, e.g. 'post|boks' ('P.O.|box'). The fullform list contains information about argument structure for some verbs. The argument structure codes used are explained in the file "norsk_ordbank_argstr.txt". Please note that this is a dump of the database athe state it was in on 1 February 2022. The latest version (1 February 2022) contains 117,445 lemmas.

新挪威语词汇数据库(Norsk Ordbank - Nynorsk)是一款面向新挪威语(Norwegian Nynorsk)的词汇数据库,其内容覆盖2012年8月1日起正式生效的官方正字法改革,以及后续针对新挪威语的拼写调整。 该数据库由基础词汇(词元,lemmas)与一组屈折变化范式(inflectional patterns)构成。每个词元关联一个或多个屈折变化范式;每个屈折变化范式包含若干条目,用于完整呈现该词元的所有屈折形式。每条目包含转换模式、词类信息与形态特征,该模式阐释了基础词如何衍生为目标屈折形式。 数据存储于七个数据表中。名为"lemma"的数据表收录了官方新挪威语词典《Nynorskordboka》的全部条目,并标注了冠词编号。全形式列表收录了所有条目依据现行官方正字法生成的全部可能屈折形式。(注:该表同时包含过度生成的形式,例如名词“snø”(雪)的推定复数形式“snøar”(雪,复数)。) 数据表"lemma_paradigme"、"paradigme"、"paradigme_boying"、"boyingsgruppe"与"boying"存储了基于基础词汇("lemma")生成全形式所需的全部信息。换言之,这些数据表搭建了词元与屈折变化范式、规则以及范畴信息之间的关联桥梁。 数据表"leddanalyse"收录了复合词的分解信息。在《Nynorskordboka》中,复合词的分解以竖线标注,例如“post|boks”对应“邮政|信箱”(原标注为'P.O.|box')。 全形式列表还包含部分动词的论元结构信息,本次所用的论元结构代码可在文件"norsk_ordbank_argstr.txt"中查阅详细说明。 请注意,本数据集为该数据库在2022年2月1日的状态快照。该版本(2022年2月1日版)共包含117445个词元。

二维码
社区交流群
二维码
科研交流群
商业服务