遇见数据集

OCR Transcriptions of the Internet Archive Early Modern Corpus (olmOCR-2)

收藏
Zenodo2026-05-28 更新2026-05-29 收录
官方服务:

资源简介:

Plain-text OCR transcriptions of 15,799 public-domain Internet Archive documents (printed works, 1614-1810), produced by olmOCR-2, used as the Chapter 3 corpus in the Master's thesis Loud Yet Invisible: A Humanist-Designed Pipeline for Unlocking the Early Modern Archive (Jacob Polay, University of Saskatchewan, 2026). Internet Archive sources only; EEBO-derived OCR is withheld pending copyright review. A manifest gives the archive.org citation for each file.

提供机构:
Zenodo
创建时间:
2026-05-28
二维码
社区交流群
二维码
科研交流群
商业服务