遇见数据集

A survey on the dataset, techniques, and evaluation metric used for abstractive text summarization

收藏
Zenodo2024-06-24 更新2024-06-25 收录
数据链接:
官方服务:

资源简介:

Whenever there is too much information out there, it is desirable to summarize. If humans are trying to create the summary, it will take lot of time. Now to make the problem of summarizing information easier and more effortless one can automate the summarization process which can reduce the time taken in creating summary. This is called as automatic summarization. The two ways of summarization are extractive summarization and abstractive summarization. Extractive summarization and its applications have been the subject of extensive research and have received state of art solution. But abstractive summarization still is a progressive field as it is difficult to create abstractive summary as humans do. Also, it is still a question i.e., how to evaluate the quality of a summary? Therefore, this paper is a comprehensive survey on the dataset used with its details and statistics, analysis of various abstractive summarization techniques and important parameters for evaluating the quality of summary. Deep leaning based models have given new direction in this field. The author also focuses on problems and challenges faced in the generation of summary which are opening the future research scope in this domain

当外界信息过载时,摘要生成便成为刚需。若由人工完成摘要撰写,则需耗费大量时间。为简化信息摘要的生成流程、降低人力投入,可通过自动化手段实现摘要生成,从而大幅缩短摘要撰写所需时长,该方法即为自动摘要(automatic summarization)。摘要生成主要分为抽取式摘要(extractive summarization)与生成式摘要(abstractive summarization)两类。抽取式摘要及其相关应用已得到广泛研究,并形成了前沿成熟的技术方案。但生成式摘要仍处于发展阶段,因其需像人类一样创作原创性摘要,实现难度较高。此外,如何科学评估摘要质量,至今仍是一个有待解决的核心问题。有鉴于此,本文针对摘要生成领域所使用的数据集及其细节与统计信息展开全面综述,同时分析了各类生成式摘要技术,并探讨了影响摘要质量的关键评估参数。基于深度学习(Deep Learning)的模型为该领域开辟了全新的发展方向。作者同时聚焦于摘要生成过程中面临的各类问题与挑战,为该领域的未来研究指明了方向与拓展空间。

创建时间:
2024-06-24
二维码
社区交流群
二维码
科研交流群
商业服务