Topic modeling private anthologies of poetry
收藏资源简介:
This digital repository complements an article on private poetry anthologies to be published in 2026 by the university presses of the Philosophical Faculty of Zagreb (FF Open Press). It contains the experimental corpus (8 private poetry anthologies, produced between 1850 and 1940, in Spanish) as well as the control corpus (54 Spanish and Latin-American poetic works in public domain), both lemmatized (Stanza) and unlemmatized, in separate .zip files. It also includes the pipeline that executed the topic modeling in Python (with comments in Spanish), the list of stopwords, and the files with the results (.csv). N.B.: In the corpus files titles, the shorthand ’COC’ (‘cancionero ordinario contemporáneo’) identifies the private anthologies. Some of the texts in the control corpus were obtained from libraries such as Project Gutenberg. All literary texts have been manually curated.



