What Data Science Taught Me About a 2,000-Year-Old Emperor's Diary
收藏资源简介:
This dataset contains the data associated with the article "What Data Science Taught Me About a 2,000-Year-Old Emperor's Diary" published on Medium. The deposit includes Python scripts and processed data files designed to extract and analyze the vocabulary of Marcus Aurelius’s Meditations. By integrating the LAGT (Latin and Ancient Greek Treebank) corpus with the LSJ (Liddell-Scott-Jones) Greek-English Lexicon, the pipeline calculates lexical variety metrics (such as Hapax Legomena percentages) and tabulates word frequencies across the traditional Stoic tripartite division of Physics, Ethics, and Logic. Included files:- Processing scripts (.py) for data extraction, filtering, and analysis.- Requirement files for environment setup.- README.txt with execution instructions. This resource is intended to support the transparency of the article's findings and provide a framework for further computational linguistic study of Stoic texts.



