遇见数据集

Dysphoria definitions generated by a large language model: bilingual corpus (Ukrainian/English), coding matrix and analysis scripts

收藏
Zenodo2026-07-13 更新2026-08-02 收录
官方服务:

资源简介:

Supplementary materials for the study "The dual role of a large language model in content analysis: source of data and instrument of its interpretation". Fifty-seven operators submitted a standardised prompt to ChatGPT in Ukrainian and English, in separate conversations and from separate accounts, yielding a paired corpus of 114 definitions of dysphoria. Texts were coded across 21 categories, with a large language model as the primary coder and blind human validation of a stratified subsample (15 pairs, approximately 26% of the corpus). The package contains the full corpus, the coding matrix, the human validation subsample, per-category Cohen's kappa computed separately by language, the codebook, and Python scripts reproducing all reported analyses (Cohen's kappa, exact McNemar test for paired cross-linguistic comparison, and Jaccard clustering on 5-grams). Operators are identified by code only; no personal data are included.

提供机构:
Zenodo
创建时间:
2026-07-13
二维码
社区交流群
二维码
科研交流群
商业服务