Institutional Newspapers Dataset
收藏资源简介:
Institutional Newspapers Dataset由哈佛法学院·Institutional Data Initiative与波士顿公共图书馆联合创建,旨在从历史报纸扫描中提取高质量结构化数据。该数据集包含从1795年至1930年的1,473,635份公共领域报纸扫描中分割出的8310万独立裁剪区域,其OCR输出总计163亿个o200k_base词元。数据集通过模块化管道生成,依次执行版面分割、OCR、类型分类、阅读顺序检测、命名实体识别、主题分类及嵌入生成等步骤,每个环节均保持可解释性与可定制性。该数据集主要用于赋能历史计算研究、改善大语言模型的预训练数据质量,并促进图书馆对馆藏报纸的深度访问与挖掘。
The Institutional Newspapers Dataset was co-created by the Harvard Law School Institutional Data Initiative and the Boston Public Library, with the objective of extracting high-quality structured data from historical newspaper scans. This dataset comprises 83.1 million independent cropped regions segmented from 1,473,635 public-domain newspaper scans spanning the period from 1795 to 1930, and its accumulated OCR outputs total 1.63 billion o200k_base tokens. The dataset is generated via a modular pipeline that sequentially performs layout segmentation, OCR, type classification, reading order detection, named entity recognition, topic classification, and embedding generation, with each step maintaining interpretability and customizability. This dataset is primarily utilized to empower historical computational research, improve the quality of pre-training data for large language models (LLMs), and facilitate in-depth access and mining of library-preserved newspaper collections.

- 1Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers哈佛法学院·Institutional Data Initiative; 波士顿公共图书馆; 哈佛法学院图书馆; 哈佛法学院·工程与应用科学学院·肯尼迪学院 · 2026年



