USB
收藏资源简介:
USB数据集是由卡内基梅隆大学和东北大学合作创建,包含1988个从维基百科中提取的文档,用于支持8种不同的文本摘要任务。数据集涵盖6个领域,包括传记、公司、学校、报纸、地标和灾难。创建过程中,研究人员首先从维基百科下载英文文章,然后使用Wikiextractor工具提取文章,去除表格和列表,保留部分头信息。数据集主要用于训练和评估模型在提取证据、纠正事实错误和生成特定主题摘要等方面的能力,旨在解决文本摘要中的关键问题,如事实正确性和可控性。
The USB Dataset was co-created by Carnegie Mellon University and Northeastern University. It contains 1988 documents extracted from Wikipedia to support 8 distinct text summarization tasks. The dataset covers 6 domains including biography, companies, schools, newspapers, landmarks and disasters. During the creation process, researchers first downloaded English articles from Wikipedia, then used the Wikiextractor tool to extract the articles, removing tables and lists while retaining partial header information. This dataset is primarily used to train and evaluate models' capabilities in evidence extraction, factual error correction and targeted topic summary generation, and aims to address key challenges in text summarization such as factual correctness and controllability.

- 1USB: A Unified Summarization Benchmark Across Tasks and Domains卡内基梅隆大学 · 2023年



