Dataset and code for "Lost in the Titles: Text-mining Metadata in the Digital Edition of Grundtvig's Works"
收藏资源简介:
This dataset accompanies the paper "Lost in the Titles: Text-mining Metadata in the Digital Edition of Grundtvig's Works" (Vad and Baunvig, 2026 fortcoming). The CSV file contains 660 titles extracted from version 1.26 of Grundtvig's Works (published 1 December 2025), the ongoing digital scholarly edition at grundtvigsværker.dk. Each row represents one published work and includes the following fields: title (main or part), year of publication, main genre, subgenre, and internal XML references from the TEI-encoded edition. The Python notebook documents the analytical workflow: title length analysis, TF-IDF word-frequency distributions, and a rule-based typological classification. The notebook also includes the LLM-assisted keyword patterns. All data derive from the CC0-licensed Grundtvig's Works edition.



