WiC
收藏资源简介:
WiC数据集是由剑桥大学等机构创建,专注于评估上下文敏感词义表示的大型数据集。该数据集包含7428个实例,每个实例包含一个目标词及其在两个不同上下文中的使用情况,旨在通过二元分类任务评估模型对词义动态变化的理解能力。数据集内容来源于WordNet、VerbNet和Wiktionary等权威资源,经过专家精心注释和筛选,确保数据质量。WiC数据集的应用领域广泛,主要用于评估和改进自然语言处理中词义消歧和上下文敏感词向量模型,以解决现有模型在处理多义词时的局限性。
The WiC dataset was developed by institutions including the University of Cambridge, and it is a large-scale benchmark dataset dedicated to evaluating context-sensitive word meaning representations. It comprises 7,428 instances, each containing a target word and its respective usages in two distinct contexts, with the goal of assessing a model's capability to comprehend dynamic shifts in word meaning via a binary classification task. The dataset's content is sourced from authoritative resources such as WordNet, VerbNet, and Wiktionary, and has been meticulously annotated and filtered by domain experts to guarantee high data quality. The WiC dataset has broad application domains, primarily used to evaluate and enhance word sense disambiguation and context-sensitive word vector models in natural language processing, thereby addressing the limitations of existing models when handling polysemous words.

- 1WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations剑桥大学 · 2019年



