遇见数据集

Data and RGF-predictions.

收藏
NIAID Data Ecosystem2026-03-08 收录
官方服务:

资源简介:

Two Chinese novels are used as empirical data i.e. A Q Zheng Zhuan (AQ for short) written by Xun Lu and Ping Fan De Shi Jie (PF for short) by Yao Lu. For each book we first remove punctuation marks and numbers from the texts, then count the Chinese characters one by one and finally get the characters frequency results. In Chinese language the words are not separated by spaces, so we use a word segmenter, Jieba (https://github.com/fxsjy/jieba), to extract words from Chinese texts. The RGF-prediction is given in the form P(k) = A′ exp(−bk)/kγ. This means that the RGF-theory transforms the data-triple (M, N, kmax) into the prediction triple (γ, b, A′).

创建时间:
2015-05-08
二维码
社区交流群
二维码
科研交流群
商业服务