遇见数据集

EarlyModernNER Training Corpus — Author's Original Annotations

收藏
Zenodo2026-05-28 更新2026-05-29 收录
官方服务:

资源简介:

Original annotations, synthetic training data, blocklists, and methodology documentation supporting the EarlyModernNER model described in Chapter 2 of the master's thesis 'Loud Yet Invisible: A Humanist-Designed Pipeline for Unlocking the Early Modern Archive' (Jacob Polay, University of Saskatchewan, 2026). Contains a 100-document gold standard hand-annotated for COMMODITY, ORGANIZATION, PERSON, and TOPONYM entities (drawn from Internet Archive sources, ~1614–1810); five versions of the annotation pipeline; synthetic positive/negative examples per entity type; hand-curated entity blocklists and whitelists; and the annotation guidelines and metrics documentation. This deposit excludes text derived from the Old Bailey Online, PCEEC2, and Royal Society Corpus (cited in the README but not redistributed); to reconstruct the full ~25-million-token training set the model was fine-tuned on, obtain those corpora separately and use the data-preparation script in the GitHub repository. See README.md for full details. License: CC-BY-4.0.

提供机构:
Zenodo
创建时间:
2026-05-28
二维码
社区交流群
二维码
科研交流群
商业服务