Replication Data for: Studying lexical dynamics and language change via generalized entropies – the problem of sample size
收藏资源简介:
Recently, it was demonstrated that generalized entropies of order α offer novel and important opportunities to quantify the similarity of symbol sequences. For the analysis of the statistical properties of natural languages, this is especially interesting since textual data are characterized by Zipf’s law, i.e. there are very few word types that occur very often (e.g. function words expressing grammatical relationships) and very many word types with a very low frequency (e.g. content words carrying most of the meaning of a sentence). Varying α makes it possible to magnify differences between different texts at specific scales of the corresponding word frequency spectrum. Here, this approach is systematically and empirically studied by analyzing the lexical dynamics of the German weekly news magazine “Der Spiegel” (consisting of approximately 365k articles and 237M words that were published between 1947 and 2017). We show that, analogous to most other measures in quantitative linguistics, similarity measures based on generalized entropies depend heavily on the sample size (i.e. text length). We argue that this makes it difficult to quantify lexical dynamics and language change and show that standard sampling approaches do not solve this problem. We discuss the consequences of the results for the statistical analysis of languages.
近期已有研究证实,阶数为α的广义熵(generalized entropies of order α)为量化符号序列的相似性提供了全新且重要的可行途径。针对自然语言的统计特性分析而言,该方法具有独特的研究价值,因为文本数据普遍遵循齐普夫定律(Zipf’s law):即仅有极少数词型(word type)会高频出现(例如用于表达语法关系的功能词(function words)),而绝大多数词型的使用频率极低(例如承载句子核心语义的实义词(content words))。通过调整α的取值,可在对应词频谱(word frequency spectrum)的特定尺度下放大不同文本间的差异。本文针对该方法展开系统性实证研究,分析对象为1947年至2017年间出版的德国新闻周刊《明镜》(Der Spiegel),该刊总计发表约36.5万篇文章、累计2.37亿词。研究表明,与定量语言学(quantitative linguistics)中的多数其他度量方法类似,基于广义熵的相似性度量对样本量(即文本长度)具有极强的依赖性。本文指出,这一特性会给词汇动态性(lexical dynamics)与语言演化的量化研究带来阻碍,同时证实标准采样方法(standard sampling approaches)无法解决该问题。最后讨论了上述研究结论对语言统计分析的相关启示与影响。



