Long term email communication data within an organization
收藏资源简介:
Undirected email correspondence between users of a large organization with over 1,000 individuals for four consecutive years (2007-2010). For this period, we have information of the sender, the receiver and the total amount of emails sent within the organization using the corporate email address. To preserve users' privacy, individuals are completely anonymized and we do not have access to email content (see Ethics statement). The data is in the following format: user1ID user2ID #emails Where #emails is the total amount of emails exchanged (sent and received) in one natural year. The files are separated by years. Ethics statement: This data is exempt from IRB review because: i) The research involves the study of existing data--email logs from 2007 to 2010, which the IT service of the organization archived routinely, as mandated by law; ii) The information is recorded by the investigators in such a manner that subjects cannot be identified, directly or through identifiers linked to the subjects. Indeed, subjects were assigned a "hash" by the IT service prior to the start of our research, so that none of the investigators can link the "hash" back to the subject. We have no demographic information of any kind, so de-anonymization is also impossible.
本数据集收录某员工规模超1000人的大型机构在2007至2010年连续四年间的无向内部邮件往来数据。在此四年周期内,数据集包含通过该机构企业邮箱发送的内部邮件的发件人、收件人及总邮件量信息。为保护用户隐私,所有个体均已完成完全匿名化处理,且研究人员无法获取邮件正文内容(详见伦理声明)。 数据集格式如下:user1ID user2ID #emails,其中#emails代表单自然年内双方往来的邮件总数量(含发送与接收)。数据文件按年份拆分存储。 伦理声明:本数据集免于伦理审查委员会(Institutional Review Board, IRB)审查,原因如下: 其一,本研究仅针对现有存档数据展开分析——2007至2010年的邮件日志由该机构信息技术部门依法依规定期归档留存; 其二,研究人员对数据的记录方式已确保无法直接或通过关联标识识别研究对象。事实上,在本研究启动前,该机构信息技术部门已为所有研究对象分配哈希值(hash),因此所有研究人员均无法将该哈希值反向关联至具体研究个体。此外,本数据集未包含任何人口统计学相关信息,故亦无法通过其他途径实现去匿名化。



