Psychological profiling from digital traces: A case study on AI-driven longitudinal analysis of personal email communications
收藏资源简介:
This is a replication data for the paper titled "Psychological profiling from digital traces: A case study on AI-driven longitudinal analysis of personal email communications" submitted for a blind review. Abstract The rapid advancement of generative Artificial Intelligence (AI) has significantly expanded opportunities for psychological research by enabling automated analysis of digital communication. This paper introduces a novel, fully automated methodology for applying Large Language Models (LLMs) to psychological text analysis, ensuring rigor through internal consistency testing, machine evaluation, and human validation. The study develops a framework for extracting psychological traits from long-form digital text and applies it across four psychological theories - Self-Determination Theory, the Big Five Personality Traits, Psychological Well-being, and Cognitive Behavioral Therapy - using a 16-year longitudinal dataset of 25,780 emails. The methodology is validated through a multi-step process, including inter-rater reliability measures and benchmarking against self-reported psychological assessments. Results confirm that LLMs can provide consistent and interpretable psychological profiling, demonstrating a structured approach that extends beyond individual-level analysis. By integrating computational psychometrics with human-computer interaction research, this study establishes a scalable method for psychological assessment from digital traces. The findings underscore the potential of generative AI to enhance behavioral research, offering a replicable framework for future studies in automated psychological analysis. The zipped file contains five csv files: Email_classification-csv: LLM (GPT-3.5 Turbo) classification of 25,780 emails for four psychological theories: SDT, Big Five, PWB and CBT. SDT_regression_data.csv Big_Five_regression_data.csv PWB_regression_data.csv CBT_regression_data.csv For 2-5 files the dependent variable is monthy percentage share of emails the were assigned a given value for categories of one of the four psychological theories analyzed. Linear regression model has been applied, where dependent variable is the percentage of emails in a specified category that assigned a specific value in this category. For example in Big Five Traits Model, for the Openness category, for each month we calculated percentage of emails that exhibit High or Low openness, or None if the content of the email does not provide enough information to assess whether the specific need is relevant. Two dependent variables were created: Openness-high and Openness-low and regressed on all independent variables. Regressions were not run for the None values. Descriptions of independent variables: - income_index: Person X salary income and consulting fees in a given month, normalized to [0,1]. - card_spending: Person X credit card expenditures in a given month, normalized to [0,1]. - abroad_far: dummy variable set to 1 for months when Person X worked in Central Asia - abroad_near: dummy variable set to 1 when Person X worked in other EU country - death_1_war: variable set to 1 in a month when Person X’ farther in law passed away. In the same month Russia invaded Ukraine. The variable was set to .75 in the following month, and to .5 in the month after that. - death_2: variable set to 1 in a month when Person X’ mother passed away. The variable was set to .75 in the following month, and to .5 in the month after that. - court_case: dummy variable set to 1 for months with the emotionally engaging inheritance court case involving other family members. - BIG4_partner: dummy variable set to 1 for months when Person X worked as a partner in BIG4 accounting firm, which resulted in adopting a professional activity sharply different from the usual Person X habits. - AI_company: dummy variable set to 1 for months when Person X worked as C-level executive at a company specializing in artificial intelligence. - elections: dummy variable set to 1 for months when Person X unsuccessfully run in parliamentary elections - covid_lockdown: dummy variable set to 1 for month where Polish government imposed tough measures during two covid lockdowns. - no_receive: number of different email recipients each month, normalized to [0,1]. - avg_length: average number of words in emails sent each month, normalized to [0,1]. While the email data was collected for January 2008 – March 2014 period, financial data was available from October 2009. There were some months where no emails with more than 10 words were sent, yielding 166 monthly observations used for regressions, before removing outliers. Independent variables were tested for multicollinearity, outlier months were removed, regressions were estimated with robust standard errors, and a range of standard tests were conducted for normality and autocorrelation of residuals, confirming good statistical properties of estimated models. Due to privacy concerns, the email texts cannot be publicly shared. However, the classifications of psychological categories derived from the email texts, along with all other relevant data, are made publicly available in this open access repository, with the consent of email author.
本数据集为提交盲审的论文《基于数字痕迹的心理画像:AI驱动的个人电子邮件通信纵向分析案例研究》的复现数据。 ## 摘要 生成式人工智能(Generative AI)的快速发展,为心理学研究开辟了全新机遇,使得自动化分析数字通信成为可能。本文提出一种新颖的全自动化方法论,将大语言模型(Large Language Model, LLM)应用于心理文本分析,并通过内部一致性检验、机器评估与人类验证确保研究严谨性。本研究构建了从长文本数字信息中提取心理特质的分析框架,并基于涵盖16年、共计25780封电子邮件的纵向数据集,结合四大心理学理论——自我决定理论(Self-Determination Theory, SDT)、大五人格特质(Big Five Personality Traits)、心理幸福感(Psychological Well-being, PWB)与认知行为疗法(Cognitive Behavioral Therapy, CBT)开展研究。 研究通过多步骤流程验证方法论有效性,包括评分者信度检验以及与自我报告心理评估的基准对比。结果证实,大语言模型可生成一致且可解释的心理画像,提出了超越个体层面分析的结构化研究路径。本研究将计算心理测量学与人机交互研究相结合,建立了一种可扩展的、基于数字痕迹的心理评估方法。研究结果凸显了生成式AI在行为研究中的应用潜力,为自动化心理分析领域的后续研究提供了可复现的分析框架。 本压缩包包含5个CSV文件: 1. Email_classification.csv:基于大语言模型(GPT-3.5 Turbo)对25780封电子邮件的分类结果,涵盖四大心理学理论(SDT、大五人格、PWB与CBT)的分类标签。 2. SDT_regression_data.csv 3. Big_Five_regression_data.csv 4. PWB_regression_data.csv 5. CBT_regression_data.csv 第2至第5个文件的因变量为:对应月度中,被归为四大心理学理论某一类别下特定取值的电子邮件占当月总邮件的百分比。 本研究采用线性回归模型,以指定类别下被赋予特定取值的电子邮件占比作为因变量。以大五人格特质模型中的开放性特质为例,我们针对每个月度计算了表现出高开放性、低开放性的电子邮件占比,若邮件内容不足以评估该特质相关性,则标记为“无”。本研究共构建两个因变量:开放性-高与开放性-低,并将其与所有自变量进行回归分析,未针对“无”标记值开展回归。 自变量说明如下: - income_index:受试者X当月薪资收入与咨询费用,经归一化处理至[0,1]区间。 - card_spending:受试者X当月信用卡消费金额,经归一化处理至[0,1]区间。 - abroad_far:虚拟变量,当月受试者X在中亚工作时取值为1。 - abroad_near:虚拟变量,当月受试者X在其他欧盟国家工作时取值为1。 - death_1_war:当月受试者X的岳父去世,且当月俄罗斯入侵乌克兰,该变量取值为1;次月取值为0.75,再次次月取值为0.5。 - death_2:当月受试者X的母亲去世,该变量取值为1;次月取值为0.75,再次次月取值为0.5。 - court_case:虚拟变量,当月受试者X涉及其他家庭成员的情感纠葛型遗产继承诉讼时取值为1。 - BIG4_partner:虚拟变量,当月受试者X在BIG4会计师事务所担任合伙人,其职业活动与受试者X日常习惯存在显著差异时取值为1。 - AI_company:虚拟变量,当月受试者X在专注于人工智能领域的企业担任首席级高管时取值为1。 - elections:虚拟变量,当月受试者X参与议会选举且落选时取值为1。 - covid_lockdown:虚拟变量,当月波兰政府针对新冠疫情实施严格封锁措施时取值为1,覆盖两次新冠封锁周期。 - no_receive:当月不同邮件收件人的数量,经归一化处理至[0,1]区间。 - avg_length:当月发送电子邮件的平均单词数,经归一化处理至[0,1]区间。 注:电子邮件数据采集时段为2008年1月至2014年3月,但财务数据仅可追溯至2009年10月。部分月份未发送单词数超过10的电子邮件,在剔除异常值前,共获得166个月度观测样本。 本研究对自变量进行了多重共线性检验,剔除了异常月度样本,采用稳健标准误开展回归估计,并针对残差的正态性与自相关性开展一系列标准检验,证实所估计模型具备良好的统计性质。 出于隐私保护考虑,电子邮件原文无法公开共享。但经邮件作者同意,本研究将从邮件文本中提取的心理类别分类结果及所有其他相关数据,通过本开放获取仓库公开提供。



