FrameNet语义框架消歧数据集
收藏资源简介:
本数据集名为FrameNet语义框架消歧数据集,由阿姆斯特丹自由大学创建,包含超过5000个来自维基百科的单词-句子对。数据集通过创新的众包方法收集,每个句子由多名工作者标注,以捕捉标注者之间的分歧。与传统方法不同,本数据集提供了一个包含基于分歧得分的框架列表,表达每个框架适用于单词的置信度。数据集的应用领域包括自然语言处理系统的训练和评估,旨在解决由于文本和框架固有的歧义导致的标注不一致问题。
This dataset, named the FrameNet Semantic Frame Disambiguation Dataset, was created by Vrije Universiteit Amsterdam and contains over 5,000 word-sentence pairs sourced from Wikipedia. The dataset was collected through an innovative crowdsourcing method, where each sentence was annotated by multiple workers to capture inter-annotator disagreement. Unlike traditional approaches, this dataset provides a list of frames paired with disagreement-based scores that reflect the confidence level of each frame being applicable to the target word. The application scope of this dataset covers the training and evaluation of natural language processing (NLP) systems, aiming to address annotation inconsistencies caused by the inherent ambiguity of both text and semantic frames.

- 1A Crowdsourced Frame Disambiguation Corpus with Ambiguity阿姆斯特丹自由大学 · 2019年



