Dysphoria definitions generated by a large language model: bilingual corpus (Ukrainian/English), coding matrix and analysis scripts
收藏资源简介:
Supplementary materials for the study "The dual role of a large language model in content analysis: source of data and instrument of its interpretation". Fifty-seven operators submitted a standardised prompt to ChatGPT in Ukrainian and English, in separate conversations and from separate accounts, yielding a paired corpus of 114 definitions of dysphoria. Texts were coded across 21 categories, with a large language model as the primary coder and blind human validation of a stratified subsample (15 pairs, approximately 26% of the corpus). The package contains the full corpus, the coding matrix, the human validation subsample, per-category Cohen's kappa computed separately by language, the codebook, and Python scripts reproducing all reported analyses (Cohen's kappa, exact McNemar test for paired cross-linguistic comparison, and Jaccard clustering on 5-grams). Operators are identified by code only; no personal data are included.



