Replication data for: Measure schematicity through information content: A quantitative approach to grammaticalization

DataONE2025-09-03 更新2025-09-13 收录

下载链接：

https://search.dataone.org/view/sha256:d407cc4d524539a0a953be94de6dbde4fe3daab70f8099d7e4dbb785839f41bf

下载链接

链接失效反馈

官方服务：

资源简介：

This is a study to propose a quantitative method to compute the schematicity of constructions, which is a key indicator of the level of grammaticalization of morphemes. In this method, to estimate the schematicity of a schema made up of two morphemes, i.e., X_ (X is the target morpheme and _ represents an open slot), we need to know the total token frequency of all types of X_, and the token frequencies of all kinds of elements occurring in the open slot. For example, if we are interested in the schematicity of “_ment”. We need to know the total token frequency of “_ment”, which is the sum of the frequencies of “shipment”, “equipment”, “employment”, “appointment” … (all types of “_ment”). We also need to know the token frequencies of “ship”, “equip”, “employ”, “appoint” … (all types of elements occurring in the open slot). Therefore, the data are morpheme bigrams (2-gram) generated from the English and Chinese corpora showing what morphemes can each morpheme combine with, together with the token frequency of each bigram, and the token frequencies of its two components respectively.

创建时间：

2025-09-04

5,000+

优质数据集

54 个

任务类型

进入经典数据集