遇见数据集

Results replication data: 16-year longitudinal analysis of emails using GPT models - testing three psychological theories

收藏
Zenodo2025-09-30 更新2026-05-26 收录
官方服务:

资源简介:

This study evaluated whether email language yields psychologically meaningful signals in Self‑Determination Theory, the Big Five, and Psychological Well‑Being. It utilized eleven LLM-based models: three based on GPT 3.5 Turbo with different prompting strategies; and eight based on AI agentic structure that acts as a teacher for knowledge-distilled multilingual transformers (students). These models are applied in a single study setting of 25,780 emails (2008–2024) from a senior executive with international experience. Privacy note Raw email texts are not included. Interested researchers can request the anonymized email corpus after the paper is published by contacting the corresponding author and signing a standard NDA, per the data‑use terms described in the manuscript. File inventory 1. emails_classification_all_models.csv (row = one email) What it is: Label outputs for every email, so you can build the monthly indices used as dependent variables in the regressions. The email body is not present; only labels, metadata, and model outputs are included. Row/column counts & coverage. 25,780 emails authored between 2008-01 and 2024-03. Total columns: 156. Key columns (metadata). Date — UTC timestamp string. WordCount — raw word count of the author’s original text (min 10, mean 66.5, max 11186). Note: although the manuscript analyzes only the first 300 words per email, this WordCount column reports the full length before that analysis‑time truncation. Label families and value spaces: • Self‑Determination Theory (SDT): Competence, Autonomy, Relatedness → values: Present, Struggle, Absent (neutral), or None (cannot be determined, neutral). • Psychological Well‑Being (PWB): Autonomy, Environmental Mastery, Personal Growth, Positive Relations with Others, Purpose in Life, Self‑Acceptance → values: Enhancing, Struggling, Maintaining (neutral), Not Applicable (neutral). • Big Five: Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism → values: High, Low, or None (cannot be determined, neutral). Column organization. Three GPT‑3.5 passes per theory (r1: zero‑shot, r2: few‑shot rater, r3: few‑shot evaluator) yield 15 Big Five, 18 PWB, and 9 SDT columns (total 42). Eight fine‑tuned multilingual transformer models contribute 56 separate‑encoder columns (four backbones × 14 dimensions) and 56 shared‑encoder columns (four backbones × 14 dimensions). 2. regression_monthly_data.csv (row = one month, key = y_m) What it is: The month‑level regressors used in the paper’s longitudinal specifications. Merge on y_m with monthly outcome indices computed from the per‑email file. Index column. y_m — string “YYYY‑MM”, covering 2009-10 through 2024-03 (inclusive). Total rows: 167. Variables and intended interpretation: income_index: monthly salary + consulting income; continuous ratio‑type index, current month value divided by full period monthly average (min ≈ 0.114, max ≈ 5.625). card_spending: [0,1] min‑max normalized credit‑card spending. abroad_far: dummy 0/1; months working in Central Asia. abroad_near: dummy 0/1; months working in another EU country. death_1_war: 1.0 in the month of father‑in‑law’s death (coincided with the start of Russia’s invasion of Ukraine), then 0.75 (t+1) and 0.5 (t+2); 0 otherwise. death_2: analogous three‑month decay for mother’s death. court_case: dummy 0/1; months with the inheritance court case. Big4_partner: dummy 0/1; months working as a Big Four partner. AI_company: dummy 0/1; months serving as a C‑level executive at an AI company. elections: dummy 0/1; months of the parliamentary campaign. covid_lockdowns: dummy 0/1; strict lockdown months in Poland (CSV uses plural 'covid_lockdowns', manuscript uses 'covid_lockdown'). no_receive: [0,1]; number of unique email recipients per month, min‑max scaled. avg_length: [0,1]; average email word count per month, min‑max scaled. Note: Regression results are reported in the paper with HAC‑robust SEs and BH‑FDR correction; they quantify contemporaneous covariation rather than causal effects.

本研究旨在验证电子邮件语言能否在自我决定理论(Self‑Determination Theory, SDT)、大五人格(Big Five)以及心理幸福感(Psychological Well‑Being, PWB)中产生具有心理学意义的信号。本研究采用了11种基于大语言模型(LLM/Large Language Model)的模型:3种基于GPT 3.5 Turbo且采用不同提示策略的模型;以及8种基于AI智能体(AI Agent)架构的模型,该架构可作为教师,为经过知识蒸馏的多语言Transformer(Transformer)学生模型提供指导。上述模型均应用于一项单研究场景,包含25780封来自一位拥有国际工作经验的高级管理人员的电子邮件,时间跨度为2008年至2024年。 隐私说明 本数据集未包含原始电子邮件文本。有需求的研究人员可在论文发表后联系通讯作者并签署标准数据保密协议(NDA),依据手稿中描述的数据使用条款,申请匿名化电子邮件语料库。 文件清单 1. emails_classification_all_models.csv(每行对应一封电子邮件) 该文件内容:所有电子邮件的标签输出,可用于构建本文回归分析中作为因变量的月度指数。文件未包含邮件正文,仅包含标签、元数据以及模型输出结果。 行列数与覆盖范围:共25780封邮件,创作时间介于2008年1月至2024年3月之间。总列数为156列。 核心元数据列: • Date:UTC时间戳字符串 • WordCount:作者原始文本的原始词数(最小值为10,均值为66.5,最大值为11186)。注:尽管手稿仅分析每封邮件的前300个词,但该WordCount列统计的是分析截断前的完整文本长度。 标签家族与取值空间: • 自我决定理论(SDT):能力(Competence)、自主性(Autonomy)、关联性(Relatedness)→ 取值为:Present(存在)、Struggle(存在困难)、Absent(缺失,中性)或None(无法判定,中性)。 • 心理幸福感(PWB):自主性、环境掌控、个人成长、积极人际关系、人生目标、自我接纳 → 取值为:Enhancing(正向提升)、Struggling(存在困境)、Maintaining(维持现状,中性)或Not Applicable(不适用,中性)。 • 大五人格:开放性(Openness)、尽责性(Conscientiousness)、外倾性(Extraversion)、宜人性(Agreeableness)、神经质(Neuroticism)→ 取值为:High(高水平)、Low(低水平)或None(无法判定,中性)。 列结构:针对每个理论进行3次GPT-3.5推理(r1:零样本(zero-shot)、r2:少样本(few-shot)标注者、r3:少样本评估者),共生成15列大五人格、18列心理幸福感以及9列自我决定理论相关列(总计42列)。8个经过微调的多语言Transformer模型贡献了56列独立编码器列(4个骨干网络 × 14个维度)以及56列共享编码器列(4个骨干网络 × 14个维度)。 2. regression_monthly_data.csv(每行对应一个月,主键为y_m) 该文件内容:本文纵向回归分析所用的月度回归变量。可通过y_m字段与基于单邮件文件计算得到的月度结果指数进行合并。 索引列:y_m — 格式为"YYYY-MM"的字符串,覆盖时间范围为2009年10月至2024年3月(含首尾),共167行。 变量与含义解释: • income_index:月度薪资+咨询收入;连续比率型指数,为当月数值除以全周期月度平均值(最小值约为0.114,最大值约为5.625)。 • card_spending:[0,1]区间内的最小-最大归一化信用卡消费金额。 • abroad_far:0/1虚拟变量;表示在中亚工作的月份。 • abroad_near:0/1虚拟变量;表示在其他欧盟国家工作的月份。 • death_1_war:在岳父去世当月取值为1.0(该月恰逢俄罗斯入侵乌克兰),t+1月取值为0.75,t+2月取值为0.5;其余月份取值为0。 • death_2:类似的三个月衰减系数,对应母亲去世事件。 • court_case:0/1虚拟变量;涉及遗产继承诉讼的月份。 • Big4_partner:0/1虚拟变量;担任四大会计师事务所合伙人的月份。 • AI_company:0/1虚拟变量;在AI公司担任C级高管的月份。 • elections:0/1虚拟变量;议会竞选活动所在的月份。 • covid_lockdowns:0/1虚拟变量;波兰实施严格封锁的月份(数据集使用复数形式covid_lockdowns,手稿中使用单数形式covid_lockdown)。 • no_receive:[0,1]区间内的数值;月度唯一邮件收件人数量,经最小-最大缩放处理。 • avg_length:[0,1]区间内的数值;月度平均电子邮件词数,经最小-最大缩放处理。 注:本文报告的回归结果采用了HAC稳健标准误与BH-FDR校正;该结果仅量化了同期共变关系,而非因果效应。

提供机构:
Zenodo
创建时间:
2025-09-30
二维码
社区交流群
二维码
科研交流群
商业服务