QAMR (Question-Answer Meaning Representation Dataset)
收藏资源简介:
问答意义表示是表示谓词-论元结构的一种新范式,它利用自由形式的问题及其答案来表示广泛的语义现象。 QAMR 的语义表达能力与现有的形式主义相比(在某些情况下超过),而表示可以由非专家注释(特别是使用众包)。 我们定义了 QAMR,开发了一个众包管道,用于在 Mechanical Turk 上大规模收集它们,收集大约 5,000 个带注释的句子的数据集,对这些数据进行彻底分析,并在数据集上运行一些内在基线。除此之外,监督开放信息提取(代码)的并行工作利用 QAMR 数据集来提高开放 IE 系统的性能,尤其是在不太常见的注释谓词(如名词化)上。
Question-Answer Meaning Representation (QAMR) is a novel paradigm for representing predicate-argument structure, which leverages free-form questions and their corresponding answers to model a wide range of semantic phenomena. The semantic expressiveness of QAMR is comparable to, and in some cases exceeds, that of existing formalisms, and its annotations can be completed by non-experts, especially via crowdsourcing. We formally define QAMR, develop a crowdsourcing pipeline for large-scale data collection on Amazon Mechanical Turk, construct a dataset containing approximately 5,000 annotated sentences, conduct thorough analysis of the collected data, and run several intrinsic baselines on the dataset. In addition, a parallel work on supervised open information extraction (code) leverages the QAMR dataset to improve the performance of open IE systems, particularly on less frequently annotated predicates such as nominalizations.




