NewSHead
收藏资源简介:
NewSHead数据集是一个多文档新闻标题数据集,用于NHNet训练新闻故事标题生成模型。该数据集包含369,940个英文新闻故事,每个故事至少包含三篇至五篇文章,用于训练、验证和测试。数据集收集自2018年5月至2019年5月的新闻文章,使用专有聚类算法根据内容相似性对文章进行分组,并由众包平台的策展人提供最多35个字符的标题来描述故事的主要信息。
The NewSHead dataset is a multi-document news headline dataset designed for training news story headline generation models using NHNet. It comprises 369,940 English news stories, each containing at least three to five articles for training, validation, and testing purposes. The dataset was compiled from news articles published between May 2018 and May 2019. Articles were grouped based on content similarity using a proprietary clustering algorithm, and curators from a crowdsourcing platform provided headlines of up to 35 characters to encapsulate the main information of each story.
数据集概述
数据集名称
NewSHead
数据集用途
用于新闻故事标题生成任务。
数据集内容
- 语言:英语
- 故事数量:369,940个新闻故事
- URL数量:932,571个唯一URL
- 数据分割:训练集359,940个故事,验证集5,000个故事,测试集5,000个故事
- 每个故事的文章数量:至少三篇,最多五篇
数据收集时间
2018年5月至2019年5月
数据处理方法
- 使用专有聚类算法,根据内容相似性对发布在时间窗口内的文章进行分组。
- 从每个聚类中选择最多五篇代表性文章用于生成故事标题。
- 通过众包平台请编辑提供最多35个字符的标题,描述故事的主要信息。
数据集下载链接
数据集处理工具
数据集引用
@InProceedings{headline2020, title = {{Generating Representative Headlines for News Stories}}, author = {Gu, Xiaotao and Mao, Yuning and Han, Jiawei and Liu, Jialu and Yu, Hongkun and Wu, You and Yu, Cong and Finnie, Daniel and Zhai, Jiaqi and Zukoski, Nicholas}, booktitle = {Proc. of the the Web Conf. 2020}, year = {2020} }




