遇见数据集

Data and code for "From Corpus to Database: An Auditable, Validated Protocol for LLM-Assisted Inductive Analysis of Dispersed Themes in Historical Periodicals" (I Diritti della Scuola, 1899-1940)

收藏
Zenodo2026-07-04 更新2026-08-01 收录
官方服务:

资源简介:

Data-and-code package accompanying the manuscript "From Corpus to Database: An Auditable, Validated Protocol for LLM-Assisted Inductive Analysis of Dispersed Themes in Historical Periodicals", under double-anonymous review at Computational Humanities Research (Cambridge University Press). Anonymised version for double-anonymous peer review. The package documents an auditable three-phase protocol (deterministic controlled-vocabulary retrieval, LLM-assisted relevance filtering, inductive multi-strategy coding) applied to the theme of clothing in relation to the teaching profession across thirty-five digitised volumes of the Italian teachers' magazine I Diritti della Scuola (1899-1940), about seventy million words. Contents: the relational SQLite database of 7,395 coded passages with the model's written justification for every judgement (plus flat CSV exports); the thematic codebook (15 dimensions with inclusion/exclusion criteria and stability ratings); the Phase 1 controlled vocabulary; the archived verbatim prompts (Italian) for the axial-coding, arbiter, and systematic-coding stages; the pipeline, validation, and figure scripts; the blind expert annotation sheets (gold set n=200, second-annotator subsample n=50, page-sample recall reading) as given to the annotators and as completed; and a script (recount_key_numbers.py) that recomputes every headline number in the paper from the database. The page images and OCR transcriptions of the source periodical are held by the Emeroteca Digitale of the Biblioteca Nazionale Centrale di Roma and are not redistributed; the source volumes and digitisation procedure are identified so the corpus can be reconstructed from the same originals.

提供机构:
Zenodo
创建时间:
2026-07-04
二维码
社区交流群
二维码
科研交流群
商业服务