SQuALITY
收藏资源简介:
SQuALITY是由纽约大学创建的一个专注于长文档摘要生成的数据集。该数据集通过聘请高素质的合同工阅读故事并从零开始编写原创摘要来构建,每个文档收集五个摘要,第一个提供概览,后续四个针对特定问题。SQuALITY基于与多选题数据集QuALITY相同的公共领域短故事构建,旨在为长上下文文本生成模型提供挑战性基准。数据集包含100个故事,500个问题,以及2000个摘要,适用于问题聚焦的抽象摘要任务,旨在解决现有自动评估指标在质量评估上的不足。
SQuALITY is a dataset focused on long-document summarization developed by New York University. It is constructed by hiring highly qualified contract workers to read short stories and write original summaries from scratch. For each story, five summaries are collected: the first provides a general overview, while the remaining four target specific questions. Built upon the same public-domain short stories as the multiple-choice dataset QuALITY, SQuALITY aims to provide a challenging benchmark for long-context text generation models. The dataset contains 100 stories, 500 questions and 2000 summaries, which is suitable for question-focused abstract summarization tasks, and is designed to address the shortcomings of existing automatic evaluation metrics in quality assessment.

- 1SQuALITY: Building a Long-Document Summarization Dataset the Hard Way纽约大学 · 2022年



