De-identified replication materials for a single-case longitudinal email-language study (2008–2024) using LLM prompting and distilled multilingual transformer students
收藏资源简介:
This Zenodo record contains a de-identified results replication package for a subject-specific (single-case) longitudinal study of personality-related language in a predominantly professional email archive spanning 2008–2024. The analysis uses theory-guided LLM-based coding (prompted GPT-3.5) and a teacher–student pipeline (GPT-4.1 agentic teacher distilled into compact multilingual transformer students) to produce email-level classifications that are aggregated into monthly indices and analyzed with HAC-robust regressions against pre-specified contextual covariates. The package includes: - Processed, de-identified email-level classification outputs (no raw email text, no subjects, no addresses; dates reduced to month start for aggregation) - Derived monthly indices, coverage summaries, and regression-ready covariate tables - Regression outputs (coefficient tables, fit summaries) and figure-generation scripts - Precomputed figures matching the manuscript (only) Not included: - Raw emails or verbatim email text - Email subjects, sender/recipient fields, or other direct identifiers - Full prompt dumps or long implementation logs To reproduce key results, unzip all ZIP files into the same directory (they unpack directly into `Data/`, `Results/`, and a small set of top-level scripts), then run `uv sync` followed by `uv run python generate_manuscript_figures.py` (see `Replication_package/replication_README.md`).



