ProRetrieval
收藏资源简介:
ProRetrieval数据集是由华东师范大学、纽约大学等机构联合构建的混合检索基准,旨在评估结合结构化过滤与语义检索的检索系统。该数据集包含两个子集:电商数据集基于Amazon ESCI和Reviews,涵盖3000个商品的结构化字段、文本和图像;邮件数据集基于Enron语料,包含5000封邮件的结构化字段和文本。每个子集由20000条训练查询和3000条测试查询组成,查询按复杂度分为单条件、二至三条件和四至五条件(含否定与嵌套)三个层级。数据集通过自动化管道生成,从候选集中采样条件并编译为可执行的混合DSL程序,再经LLM改写为自然语言查询,旨在解决真实场景中需同时处理结构化约束与语义意图的复杂检索问题,为混合检索系统的训练与评估提供标准化基准。
The ProRetrieval dataset is a hybrid retrieval benchmark jointly constructed by institutions including East China Normal University, New York University and other organizations. It is designed to evaluate retrieval systems that combine structured filtering and semantic retrieval. This dataset contains two subsets: the e-commerce dataset, based on Amazon ESCI and Reviews, covers structured fields, text and images of 3000 products; the email dataset, based on the Enron corpus, contains structured fields and text of 5000 emails. Each subset consists of 20,000 training queries and 3,000 test queries, which are divided into three complexity levels: single-condition, two-to-three-condition, and four-to-five-condition (including negation and nesting). The dataset is generated through an automated pipeline: conditions are sampled from candidate sets and compiled into executable hybrid DSL programs, which are then rewritten into natural language queries by LLMs. It aims to solve complex retrieval problems that require simultaneous processing of structured constraints and semantic intentions in real-world scenarios, providing a standardized benchmark for the training and evaluation of hybrid retrieval systems.

- 1ProRetrieval: Learning to Orchestrate Hybrid Search via Executable Program Synthesis华东师范大学; 纽约大学; Matter Innovation Inc.; ThinRedLine; 山东科技大学; 暨南大学; 独立研究者; 复旦大学 · 2026年




