遇见数据集

A Ground-Truth Dataset for Article Separation in Historical Newspapers: A ProQuest corpus centered on the China Institute in America (1926-1952)

收藏
Zenodo2025-12-13 更新2026-05-26 收录
官方服务:

资源简介:

Overview This ground-truth dataset contains manually segmented documents, with partial post-OCR correction, derived from an original corpus of news articles focused on the China Institute in America, drawn from the ProQuest collection of Chinese Historical newspapers. The dataset includes 96 articles published between 1926 and 1952. The ground-truth data contains the following fields: DocId: Unique identifier as stored in the Modern China Textual Database (MCTB). Date: Original date of publication Title: Article title as provided by ProQuest. Source: Periodical in which the article was published (principally China Press, China Weekly Review, North-China Herald) Text: Original, unsegmented text text_seg: Historian-curated segmented text, produced using GPT + close reading length: Character/word length of the original text length_seg: Character/word length after re-segmentation diff: length difference between original and segmented text The segmentation process uses a hybrid human–AI workflow: an automated step with a GPT-based “Historical Text Segmenter,” followed by detailed historian-guided verification and correction. The result is a high-quality ground-truth dataset suitable for OCR benchmarking, segmentation modeling, historical text analysis, and digital humanities research. Additional documentation on the configuration of the GPT “Historical Text Segmenter” is available here. Use Cases This dataset is intended for: Historical research on Sino-American cultural institutions Media and discourse analysis of Shenbao Training/evaluating segmentation and OCR models Digital humanities projects requiring high-quality ground truth corpora Studies of textual reuse and viral news circulation in Republican-era newspapers

数据集概述 本真值(ground-truth)数据集包含经人工分段且已完成部分光学字符识别(Optical Character Recognition,OCR)后校正的文档,其源自以美国中国研究所(China Institute in America)为主题的新闻文章原始语料库,该语料库取自ProQuest馆藏的中文历史报纸。本数据集涵盖1926年至1952年间发表的96篇文章。 该真值数据集包含以下字段: DocId:存储于近代中国文本数据库(Modern China Textual Database,MCTB)的唯一标识符 Date:文章的原始发表日期 Title:ProQuest提供的文章标题 Source:文章发表的期刊(主要包括《中国报》《中国周刊》《北华捷报》) Text:未经分段的原始文本 text_seg:经历史学家审定的分段文本,通过GPT结合细读分析生成 length:原始文本的字符/单词数量 length_seg:重新分段后的文本字符/单词数量 diff:原始文本与分段后文本的长度差值 本次分段流程采用人机协同混合工作流:首先依托基于GPT的「历史文本分段器」完成自动化分段,随后由历史学家进行细致的人工校验与修正。最终产出的高质量真值数据集,可用于OCR基准测试、分段模型训练、历史文本分析以及数字人文研究。有关该GPT「历史文本分段器」的配置详情可查阅此处附加文档。 应用场景 本数据集适用于以下场景: 1. 美中文化机构相关历史研究 2. 《申报》媒体与话语分析 3. 分段模型与OCR模型的训练与评估 4. 需要高质量真值语料库的数字人文项目 5. 民国时期报纸文本复用与病毒式新闻传播研究

提供机构:
Zenodo
创建时间:
2025-12-13
二维码
社区交流群
二维码
科研交流群
商业服务