MULTI-NEWS+
收藏资源简介:
MULTI-NEWS+是由韩国中央大学的研究团队开发的一个用于多文档摘要任务的数据集。该数据集包含56,216个文档集合,每个集合由新闻文章组成,旨在通过清洗策略提高现有数据集的质量。数据集创建过程中,研究团队利用大型语言模型(LLMs)进行数据标注,通过链式思维(CoT)和多数投票等方法模仿人类标注,有效识别并排除与摘要无关的文档。MULTI-NEWS+的应用领域主要集中在自然语言处理中的多文档摘要任务,旨在通过高质量的数据集提升模型性能和可靠性。
MULTI-NEWS+ is a multi-document summarization dataset developed by a research team at Chung-Ang University in the Republic of Korea. Comprising 56,216 document collections each made up of news articles, this dataset aims to elevate the quality of existing datasets through targeted cleaning strategies. During its development, the research team employed Large Language Models (LLMs) for data annotation, replicating human annotation workflows via approaches including Chain-of-Thought (CoT) and majority voting to efficiently identify and filter out documents unrelated to the target summaries. Primarily applied to multi-document summarization tasks in natural language processing, MULTI-NEWS+ is designed to boost model performance and reliability by providing high-quality training data.




