EarlyModernNER Training Corpus — Author's Original Annotations
收藏资源简介:
Original annotations, synthetic training data, blocklists, and methodology documentation supporting the EarlyModernNER model described in Chapter 2 of the master's thesis 'Loud Yet Invisible: A Humanist-Designed Pipeline for Unlocking the Early Modern Archive' (Jacob Polay, University of Saskatchewan, 2026). Contains a 100-document gold standard hand-annotated for COMMODITY, ORGANIZATION, PERSON, and TOPONYM entities (drawn from Internet Archive sources, ~1614–1810); five versions of the annotation pipeline; synthetic positive/negative examples per entity type; hand-curated entity blocklists and whitelists; and the annotation guidelines and metrics documentation. This deposit excludes text derived from the Old Bailey Online, PCEEC2, and Royal Society Corpus (cited in the README but not redistributed); to reconstruct the full ~25-million-token training set the model was fine-tuned on, obtain those corpora separately and use the data-preparation script in the GitHub repository. See README.md for full details. License: CC-BY-4.0.



