Text Reuse in GEI-Digital Textbooks from Lower Saxony Detected with Ikhneutes
收藏资源简介:
This file contains a corpus of historical German textbooks annotated with text reuse detected by n-gram shingling. It was created in the project “The Evolution of the Textbook. Text reuse detection for the analysis of the influence of early textbook production in Lower Saxony on the development of the genre as a whole [TextbookEvolution]”, funded by the Lower Saxony Ministry for Science and Culture. Text reuse detection was carried out with Ikhneutes, a Rust-based pipeline and viewer for text reuse detection, schedule for release for spring/summer 2026. The corpus comprises 462 textbooks published between 1762 and 1918 in what today is Lower Saxony, Germany: mostly in Braunschweig, Göttingen, Hannover, Oldenburg, Osnabrück, and Wolfenbüttel. The subjects are geography, history, “Realien” (natural sciences), German primers, and reading books. The books were originally digitized by the Leibniz Institute for Educational Media | Georg Eckert Institute (GEI) for their digital collection of historical textbooks “GEI-Digital”; the underlying source data derive from the Projektkorpus “Schulbuch-Evolution” aus “GEI-Digital” – Annotierte Daten, that is, a selective data extraction from GEI-Digital whose TEI-XML sources were linguistically annotated by Frank Wiegand of the Zentrum für digitale Lexikographie der deutschen Sprache (ZDL), using DTA::CAB and distributed in DDC-Tabs format. Corpus statistics: Total books: 462 Total pages: 114,381 Total tokens: 50,529,431 Detected reuse instances (individual correlated passages): 106,491 (these have not yet been systematically evaluated for false positives) Books with detected reuse: 460 Book pairs with reuse relations: 33,447 The file itself is an export of a SQLite 3 database generated by Ikhneutes. It is suitable for archival distribution, SQL-based inspection, and downstream processing. The database contains document records and metadata, token and lemma occurrences with positional information, and computed text-reuse annotations between documents. It was built from project-specific TAB-based tokenized inputs following the common corpus schema. The file can be opened with any SQLite-compatible tool and reused independently of the application itself, although inspection through the Ikhneutes GUI may in many cases be more convenient.



