TROVE
收藏资源简介:
TROVE数据集是由中国科学院自动化研究所等多个机构创建的,旨在追踪目标文本中每个句子回溯到特定源句子,并标注细粒度的关系类型(如引用、压缩、推理等)。该数据集基于三个公开数据集(LongBench、LooGLE和CRUD-RAG)构建,涵盖了11种不同的应用场景,支持多文档和长文档的追踪。数据集的构建经历了句子检索、GPT-4自动标注和人工审核三个阶段,以确保高质量、细粒度的起源数据。
The TROVE dataset was developed by multiple institutions including the Institute of Automation, Chinese Academy of Sciences. It is designed to trace each sentence in a target text back to its specific source sentence, and annotate fine-grained relationship types such as citation, compression, inference, and more. Built upon three public datasets (LongBench, LooGLE, and CRUD-RAG), this dataset covers 11 distinct application scenarios and supports tracking across multi-document and long-form documents. The dataset construction involves three stages: sentence retrieval, GPT-4 automatic annotation, and manual review, to ensure the production of high-quality, fine-grained provenance data.

- 1TROVE: A Challenge for Fine-Grained Text Provenance via Source Sentence Tracing and Relationship Classification中国科学院自动化研究所 · 2025年



