QUEST
收藏资源简介:
QUEST数据集由宾夕法尼亚大学和Google DeepMind共同创建,包含3357条自然语言查询,这些查询隐含了集合操作,如交集、并集和差集。数据集挑战模型匹配查询中的多个约束与文档中的相应证据,并正确执行各种集合操作。数据集通过半自动方式构建,使用维基百科类别名称,自动从单个类别生成查询,然后通过众包工作者进行改写和进一步验证自然性和流畅性。众包工作者还评估实体的相关性,并突出显示查询约束到文档文本的归属。数据集的应用领域包括分析检索系统在处理此类查询时的性能,旨在解决检索系统在处理复杂查询时的挑战。
The QUEST dataset was co-created by the University of Pennsylvania and Google DeepMind, comprising 3357 natural language queries that implicitly involve set operations such as intersection, union, and set difference. This dataset challenges models to match multiple constraints in the queries with corresponding evidence in documents and correctly execute various set operations. The dataset was constructed via a semi-automated pipeline: it first automatically generates queries from individual Wikipedia category names, then has crowdworkers rewrite the generated queries and further validate their naturalness and fluency. Additionally, crowdworkers evaluate the relevance of entities and highlight the alignment between query constraints and document text. The application scenarios of this dataset include analyzing the performance of retrieval systems when handling such queries, with the goal of addressing the challenges encountered by retrieval systems when processing complex queries.




