遇见数据集

Code and data for “From Face to Relations: Politeness strategies in Enron’s workplace email”

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This repository contains the data and code supporting the study “From Face to Relations: Politeness Strategies in Enron’s Workplace Email.” It enables full transparency and reproducibility of the analytical pipeline, including data construction, network modeling, and statistical analysis. The folder discussion_output provides the processed datasets used to generate the figures in the article. These files correspond directly to the visualizations reported in the Discussion section and can be used to reproduce all plotted results. The files case_candidates_by_PDR (case1–4).csv and case_candidates_by_role (case5–10).csv document the selection procedure for qualitative case analyses. The former identifies candidate emails based on the P–D–R framework (Power, Distance, and Imposition), while the latter selects cases according to organizational roles within Enron (e.g., CEO and other hierarchical positions). These files ensure that all illustrative examples in the paper are traceable and systematically derived. The script discussion_experiments.py contains the full data analysis workflow used to produce the results in the Discussion section. It corresponds directly to the outputs stored in the discussion_output folder, including statistical summaries and plotting-ready data. The dataset messages_R_with_politeness.csv is the core message-level file. It includes email metadata and text-derived features, such as request identification, imposition score (R), five decomposed dimensions (cost/effort, urgency, risk/accountability, autonomy constraint, and dependency blocking), and multi-label annotations of four politeness strategies (bald-on-record, positive politeness, negative politeness, and off-record). The file node_metrics.csv contains node-level network attributes (e.g., centrality measures, clustering, community assignment, and power index), while dyadic_metrics.csv provides edge-level attributes (e.g., tie strength, relative power, and distance-related measures). The file global_metrics.txt reports overall network statistics. Together, these resources support a multi-level analysis that integrates pragmatics with social network structure, allowing other researchers to replicate, validate, and extend the findings.

本仓库包含支撑研究《从面孔到关系:安然职场邮件中的礼貌策略》(From Face to Relations: Politeness Strategies in Enron’s Workplace Email)的全部数据与代码,可实现分析流程的全透明化与可复现性,涵盖数据构建、网络建模与统计分析三大环节。 文件夹`discussion_output`存放了用于生成论文图表的经预处理数据集,这些文件与论文“讨论”章节中呈现的可视化内容一一对应,可直接用于复现所有绘图结果。 `case_candidates_by_PDR (case1–4).csv`与`case_candidates_by_role (case5–10).csv`两份文件记录了定性案例分析的筛选流程:前者基于P-D-R框架(权力、距离与强加性)筛选候选邮件,后者则依据安然公司内部的组织角色(如首席执行官与其他层级岗位)选取案例,确保论文中所有示例案例均可溯源且具备系统性推导依据。 脚本`discussion_experiments.py`包含了用于生成“讨论”章节结果的完整数据分析工作流,与`discussion_output`文件夹中存储的输出内容一一对应,涵盖统计汇总结果与可直接用于绘图的预处理数据。 数据集`messages_R_with_politeness.csv`为核心的邮件级文件,包含邮件元数据与文本衍生特征,例如请求识别结果、强加性得分(R)、五个分解维度(成本/努力、紧迫性、风险/问责性、自主性约束与依赖性阻碍),以及四种礼貌策略的多标签标注:直接显性策略(bald-on-record)、积极礼貌策略(positive politeness)、消极礼貌策略(negative politeness)与非公开策略(off-record)。 文件`node_metrics.csv`包含节点级网络属性(如中心性指标、聚类系数、社区分配与权力指数),而`dyadic_metrics.csv`提供边级属性(如纽带强度、相对权力与距离相关指标)。文件`global_metrics.txt`则报告了整体网络统计量。 综上,本仓库的全部资源支持将语用学与社会网络结构相结合的多维度分析,可供其他研究人员复现、验证并拓展本研究的发现。

创建时间:
2026-03-25
二维码
社区交流群
二维码
科研交流群
商业服务