MEDIASUM
收藏资源简介:
MEDIASUM是一个大规模的媒体采访数据集,包含463.6K条来自NPR和CNN的采访转录及其摘要。该数据集通过收集NPR和CNN的采访转录,并使用概述和主题描述作为摘要来创建。MEDIASUM不仅规模大,还包含了多领域的复杂多方对话,适用于对话摘要研究。数据集的创建过程中,对CNN的采访进行了主题匹配的分割处理,以提高数据集的质量和适用性。MEDIASUM主要用于改进对话摘要模型的性能,特别是在转移学习方面,能够提升模型在其他对话摘要任务上的表现。
MEDIASUM is a large-scale media interview dataset consisting of 463.6K interview transcripts and their corresponding summaries sourced from NPR and CNN. This dataset is constructed by collecting interview transcripts from NPR and CNN, and using their overviews and topic descriptions as the associated summaries. Beyond its large scale, MEDIASUM encompasses multi-domain complex multi-party conversations, making it highly suitable for conversational summarization research. During the dataset creation process, topic-aligned segmentation was performed on CNN interviews to improve the dataset's quality and applicability. MEDIASUM is primarily used to enhance the performance of conversational summarization models, especially in transfer learning scenarios, where it can elevate the model's performance on other conversational summarization tasks.

- 1MediaSum: A Large-scale Media Interview Dataset for Dialogue Summarization微软认知服务研究组 · 2021年



