Chinese Newspaper Similarity Dataset (1910–1945): TF-IDF cosine similarity across seven Republican-era newspapers
收藏资源简介:
Data, source provenance, figures, and analysis outputs for the study "The Evolution of Written Chinese (1910–1945): A Computational Study of Newspapers." Seven Republican-era Chinese newspapers (Shenbao 申報, Shangwu Ribao 商務日報, Beijing Ribao 北京日報, Yuehua Bao 粵華報, Hunan Guomin Ribao 湖南國民日報, Xinxin Xinwen 新新新聞, and Jiangsheng Bao 江聲報) were sourced from the Internet Archive (East Asian Newspapers and Periodicals, 1850–1950), OCR'd with two independent pipelines (Google Cloud Vision Document Text Detection, prioritizing Traditional Chinese, and Tesseract), normalized to 50,000 characters per newspaper-year, and compared using TF-IDF–weighted cosine similarity. This deposit contains the pairwise similarity tables (vision and text measures), the 1938 cross-sectional matrix, the original analysis workbooks, all summary and pairwise figures, a machine-readable provenance table for the seven source collections (58,658 issues in total; 38,013 within the 1910–1945 study window), a table of the years actually scanned per newspaper, and helper scripts to regenerate the issue-level manifest from the Internet Archive API. The raw newspaper page images are referenced by Internet Archive identifier rather than redistributed. Derived data, figures, code, and documentation are released under CC BY 4.0.



