Zipf’s Law in China’s Local Government Work Reports: A 21-year Study Using Natural Language Processing and Regression Analysis
收藏资源简介:
This study presents the first large-scale empirical investigation of Zipf’s Law in Chinese provincial government work reports (2003-2023), utilizing a corpus of 651 reports. Employing natural language processing techniques (including Jieba word segmentation with a custom dictionary) and a double-logarithmic regression model, we analyzed word frequency distributions. Results indicate that while generally conforming to Zipf’s Law, substantial inter-regional and inter-temporal variation exists. This variation may be attributable to factors beyond the scope of this study, such as region-specific policies or the influence of the 18th National Congress of the Communist Party of China. While our findings largely confirm Zipf’s Law’s applicability to this specific corpus, the limitations of this study include potential biases in word segmentation and the exclusion of county-level reports. Future research should address these limitations by incorporating a broader range of administrative levels and conducting cross-cultural comparisons with other countries’ political documents. Further investigation of other quantitative linguistic laws (e.g., Heaps’Law, Menzerath’s Law) within this corpus is also warranted.
本研究针对2003-2023年中国省级政府工作报告,基于包含651份报告的语料库,开展了首次大规模的齐普夫定律(Zipf’s Law)实证考察。本研究采用自然语言处理技术(含搭载自定义词典的结巴分词(Jieba))与双对数回归模型,对词频分布展开分析。结果显示,尽管整体符合齐普夫定律,但研究样本存在显著的区域间与跨时间差异。此类差异可归因于本研究未涵盖的诸多因素,例如区域专属政策,或是中国共产党第十八次全国代表大会的影响。本研究结果在很大程度上证实了齐普夫定律适用于该语料库,但本研究存在一定局限:分词环节可能存在偏差,且未纳入县级政府工作报告。未来研究可通过覆盖更广泛的行政层级,并与其他国家的政治文件开展跨文化对比,以弥补上述局限。此外,针对该语料库开展希普斯定律(Heaps’Law)、门泽拉斯定律(Menzerath’s Law)等其他量化语言定律的进一步研究亦值得开展。




