SCOPE
收藏资源简介:
SCOPE数据集是由阿姆斯特丹大学与约翰斯·霍普金斯大学等机构联合构建的覆盖感知检索训练资源,旨在解决长格式检索增强生成中信息覆盖不足的挑战。该数据集包含9万条训练对,源自Researchy Questions的查询及其分解子问题,通过Llama-3 70B模型生成子问题可回答性标注来增强覆盖信号。其构建过程通过合成多角度查询和覆盖评分,专门用于训练能够同时优化相关性和信息覆盖度的检索模型,主要应用于长格式RAG系统,以提升事实覆盖的全面性。
The SCOPE dataset is a coverage-aware retrieval training resource jointly developed by the University of Amsterdam, Johns Hopkins University and other institutions, aiming to address the challenge of insufficient information coverage in long-form retrieval-augmented generation (RAG). This dataset includes 90,000 training pairs sourced from queries and their decomposed sub-questions within the Researchy Questions dataset, with its coverage signals enhanced by annotative labels for sub-question answerability generated by the Llama-3 70B model. Its construction workflow synthesizes multi-angle queries and calculates coverage scores, and is specifically designed to train retrieval models that jointly optimize both relevance and information coverage. It is primarily deployed in long-form RAG systems to improve the comprehensiveness of factual coverage.
数据集概述
数据集名称:DylanJHJ/scope
许可证:Apache-2.0
说明:该数据集页面未提供详细的描述信息、数据构成、使用方式或相关示例。目前仅公开了许可证信息,为 Apache-2.0 开源许可。

- 1Search for Coverage: Learning Coverage-Aware Retrieval with Augmented Sub-Question Answerability阿姆斯特丹大学; 约翰斯·霍普金斯大学; 莱顿大学 · 2026年



