SARA - A Collection of Sensitivity-Aware Relevance Assessments
收藏资源简介:
<strong>SARA - A Collection of Sensitivity-Aware Relevance Assessments</strong> Presented here is a collection of Sensitivity-Aware Relevance Assessments for the UC Berkely labelled subset of the Enron Email Collection. The Hearst [1] labelled version of the Enron Email Collection is a subset of the CMU collection that contains 1702 emails that were annotated as part of a class project at UC Berkley. Students in the Natural Language Processing course were tasked with annotating the emails as relevant or not relevant to 53 different categories. Therefore, the labelled version of the Enron email collection provides a rich taxonomy of labels which can be used for multiple definitions of sensitivity such as the Purely Personal and Personal but in a Professional Context. The categories that the emails are labelled for can be seen in [Table 1](#table-1). The files for the labelled version of the Enron Email Collection are available from the UC Berkely website. We deploy a topic modelling approach to identify topical themes in the labelled Enron collection that serve as a basis for our information needs which are in turn used to gather queries and relevance assessments, the notebook for which is available here. Two separate crowdsourcing tasks are carried out in the development of SARA. Firstly, query formulations are crowdsourced to represent the information needs and, secondly, relevance assessments are crowdsourced for a pooled set of documents from the labelled Enron collection for each of the information needs. The SARA Collection of Sensitivity-Aware Relevance Assessments is available through the popular ir_datasets library. More information can be found on the ir_datasets GitHub and website. <strong>Information Needs</strong> To create our set of sensitivity-aware relevance assessments for the labelled Enron email collection, we first identify a set of topical subjects that reflect the contents of the emails in the collection. We use a topic modelling approach to identify the information needs. When identifying topics to be used as information needs, we are interested in identifying general themes that relate to the topics of discussion that might likely be covered in the contents (i.e., the body) of the emails in the collection. The topics are chosen to be broad enough to be able to reasonably expect that there would be relevant documents in the collection, and not so specific that it would require specialist knowledge to make a judgement of relevance on the subject. Subsequently, we manually construct short passages of text to serve as descriptions of the information needs that are to be searched for in the collection by the crowdworkers. The information needs that the crowdworkers are available in the <em>information_needs.tsv </em>file. <strong>Queries</strong> In order to collect relevance assessments for pairs of emails and information needs, different query formulations are first needed to generate pools of documents. Query formulations for each topic are collected from crowdworkers from the Prolific crowdwork platform. Ten information needs are shown to each crowdworker and they are asked to provide a query formulation that they would use to get relevant documents to satisfy the information need they are presented with. Three queries for each of the fifty information needs are released. The resulting queries are available in the <em>repeated_queries.tsv</em> file. <strong>Relevance Assesments</strong> Crowdworkers are shown an information need and an email and asked to rate the document as being either <em>Highly Relevant</em>, <em>Partially Relevant</em>, or <em>Not Relevant</em> to the information need. Each information need/email pair is judged by three crowdworkers and a majority vote is used to generate a ground truth label. Since each information need / email pair is judged by three crowdworkers and there are three possible labels, it is possible for each of the labels to be selected by one crowdworker. In practice, this only happened for 134 pairs. In such cases, ties are broken by having one of the authors read the document and make an additional judgement. In order to ensure that sensitive documents definitely have relevance labels they were also judged by one of the authors for each of the information needs. The relevance assessments are available in the <em>repeated_qrels.txt</em> file. The relevance assessments are in the format 'query iteration document relevancy'. The iteration column is used for IR_Datasets and can be safely ignored and the document name is the filename used in the labelled Enron collection. <em>Table 1</em> 1) Coarse genre 2) Included/forwarded information 3) Primary topics (If coarse genre 1.1 is selected) 4) Emotional tone (If not neutral) 1.1 Company Business, Strategy, etc. (See 3) 2.1 Includes new text in addition to forwarded material 3.1 Regulations and regulators (includes price caps) 4.1 Jubilation 1.2 Purely Personal 2.2 Forwarded email(s) including replies 3.2 Internal projects -- progress and strategy 4.2 Hope / anticipation 1.3 Personal but in professional context (e.g., it was good working with you) 2.3 Business letter(s) / document(s) 3.3 Company image -- current 4.3 Humor 1.4 Logistic Arrangements (meeting scheduling, technical support, etc.) 2.4 News article(s) 3.4 Company image -- changing / influencing 4.4 Camaraderie 1.5 Employment arrangements (job seeking, hiring, recommendations, etc.) 2.5 Government / academic report(s) 3.5 Political influence / contributions / contacts 4.5 Admiration 1.6 Document editing/checking (collaboration) 2.6 Government action(s) (such as results of a hearing, etc.) 3.6 California energy crisis / California politics 4.6 Gratitude 1.7 Empty message (due to missing attachment) 2.7 Press release(s) 3.7 Internal company policy 4.7 Friendship / affection 1.8 Empty message 2.8 Legal documents (complaints, lawsuits, advice) 3.8 Internal company operations 4.8 Sympathy / support 2.9 Pointers to url(s) 3.9 Alliances / partnerships 4.9 Sarcasm 2.10 Newsletters 3.10 Legal advice 4.10 Secrecy / confidentiality 2.11 Jokes, humor (related to business) 3.11 Talking points 4.11 Worry / anxiety 2.12 Jokes, humor (unrelated to business) 3.12 Meeting minutes 4.12 Concern 2.13 Attachment(s) (assumed missing) 3.13 Trip reports 4.13 Competitiveness / aggressiveness 4.14 Triumph / gloating 4.15 Pride 4.16 Anger / agitation 4.17 Sadness / despair 4.18 Shame 4.19 Dislike / scorn The Sensitivity-Aware Relevance Assessments dataset is held under an Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) licence which allows for it to be adapted, transformed and built upon. Questions and comments are welcomed via email. <strong>References</strong> [1] Marti A Hearst. 2005. Teaching applied natural language processing: Triumphs and tribulations. In Proc. of Workshop on Effective Tools and Methodologies for Teaching NLP and CL.
<strong>SARA:敏感度感知相关性评估数据集</strong> 本数据集收录了针对加州大学伯克利分校(UC Berkeley)标注的安然邮件集合(Enron Email Collection)子集的敏感度感知相关性评估。由马蒂·A·赫斯特[1]标注的安然邮件集合是卡内基梅隆大学(CMU)集合的子集,包含1702封邮件,该标注工作是加州大学伯克利分校一门自然语言处理(Natural Language Processing)课程的班级项目内容。该课程的学生需将邮件标注为与53个不同类别相关或不相关。因此,该标注版安然邮件集合拥有丰富的标签分类体系,可用于多种敏感度定义,例如“纯私人内容”与“专业场景下的私人内容”。邮件所对应的标注类别详见[表1](#table-1)。该标注版安然邮件集合的文件可从加州大学伯克利分校官网获取。 我们采用主题建模(Topic Modelling)方法识别该标注版安然邮件集合中的主题,以此作为信息需求的基础,进而生成查询与相关性评估,对应的代码笔记可在此处获取。SARA的构建包含两项独立的众包任务:其一,众包生成查询表述以代表各类信息需求;其二,针对每项信息需求,从标注版安然邮件集合的池化文档集中开展相关性评估的众包工作。SARA敏感度感知相关性评估数据集可通过流行的ir_datasets库获取,更多信息可查看ir_datasets的GitHub仓库与官网。 <strong>信息需求</strong> 为构建针对该标注版安然邮件集合的敏感度感知相关性评估集,我们首先需识别一组可反映该邮件集合内容的主题对象。我们采用主题建模方法来确定信息需求。在选择作为信息需求的主题时,我们关注的是与邮件正文讨论话题相关的通用主题:这类主题应足够宽泛,以确保集合中存在相关文档,同时又不至于过于具体,需要专业知识才能判断相关性。随后,我们手动构建简短文本段落,作为将由众包工作者在集合中搜索的信息需求的描述。可供众包使用的信息需求存储在<em>information_needs.tsv</em>文件中。 <strong>查询</strong> 为收集邮件与信息需求对的相关性评估,首先需要不同的查询表述来生成文档池。每个主题的查询表述从Prolific众包平台的工作者处收集。每位众包工作者会收到10个信息需求,并被要求提供一个用于获取相关文档以满足该信息需求的查询。针对50个信息需求中的每一个,我们收集了3个查询。最终生成的查询存储在<em>repeated_queries.tsv</em>文件中。 <strong>相关性评估</strong> 众包工作者会看到一个信息需求与一封邮件,被要求将该文档评定为<em>高度相关</em>、<em>部分相关</em>或<em>不相关</em>。每个信息需求/邮件对会由3位众包工作者进行标注,采用多数投票来生成真实标签。由于每个信息需求/邮件对由3位工作者标注,且存在3种可能的标签,因此存在每位标签各得一票的情况,实际中这种情况仅出现于134个样本对。对于此类平局,我们会由一名作者阅读文档并给出额外标注以消解平局。为确保敏感文档都拥有相关性标签,我们还会让一名作者针对每个信息需求对这些文档进行标注。相关性评估存储在<em>repeated_qrels.txt</em>文件中,其格式为“查询 迭代 文档 相关性”。其中迭代列专为ir_datasets库设计,可安全忽略;文档名即该标注版安然邮件集中使用的文件名。 <em>表1</em> 1) 粗分类 2) 包含/转发的信息 3) 主要话题(若选择粗分类1.1) 4) 情感基调(若非中性) 1.1 公司业务、战略等(参见3) 2.1 转发材料之外包含新文本 3.1 监管与监管机构(含价格上限) 4.1 喜悦 1.2 纯私人内容 2.2 转发的邮件(含回复) 3.2 内部项目——进展与战略 4.2 期望/预期 1.3 专业场景下的私人内容(例如“很高兴与您共事”) 2.3 商业信函/文档 3.3 公司形象——当前状况 4.3 幽默 1.4 后勤安排(会议安排、技术支持等) 2.4 新闻文章 3.4 公司形象——改变/塑造 4.4 情谊 1.5 雇佣安排(求职、招聘、推荐信等) 2.5 政府/学术报告 3.5 政治影响力/捐赠/人脉 4.5 赞赏 1.6 文档编辑/审核(协作) 2.6 政府行动(例如听证会结果等) 3.6 加州能源危机/加州政治 4.6 感激 1.7 空邮件(因缺失附件) 2.7 新闻稿 3.7 公司内部政策 4.7 友谊/关爱 1.8 空邮件 2.8 法律文档(投诉、诉讼、建议) 3.8 公司内部运营 4.8 同情/支持 2.9 网址链接 3.9 联盟/合作 4.9 讽刺 2.10 时事通讯 3.10 法律建议 4.10 保密/机密 2.11 与业务相关的笑话、幽默 3.11 谈话要点 4.11 担忧/焦虑 2.12 与业务无关的笑话、幽默 3.12 会议纪要 4.12 顾虑 2.13 附件(疑似缺失) 3.13 行程报告 4.13 竞争力/攻击性 4.14 得意/幸灾乐祸 4.15 自豪 4.16 愤怒/烦躁 4.17 悲伤/绝望 4.18 羞愧 4.19 厌恶/轻蔑 本敏感度感知相关性评估数据集采用署名-非商业性使用4.0国际(CC BY-NC 4.0)许可协议发布,允许对数据集进行改编、转换和二次创作。欢迎通过邮件提出疑问与意见。 <strong>参考文献</strong> [1] Marti A Hearst. 2005. 教授应用自然语言处理:成就与困境。收录于:高效工具与方法论研讨会论文集,用于教授自然语言处理与计算语言学。



