Syn-D-CNN Dataset

Name: Syn-D-CNN Dataset
Creator: figshare
Published: 2025-05-01 06:12:48
License: 暂无描述

DataCite Commons2025-05-01 更新2024-11-06 收录

下载链接：

https://figshare.com/articles/dataset/Syn-D-CNN_Dataset/27367917/1

下载链接

链接失效反馈

官方服务：

资源简介：

Text summarization condenses extensive content into concise summaries; however, current approaches often rely on large language models (LLMs), which can lack interpretability and are susceptible to generating hallucinated content. To address these issues, we propose Docusage, an interpretable framework that replicates human summaries through a hierarchical clustering approach combined with extractive summarization, augmented by selective, LLM-based abstraction. Docusage minimizes the risk of hallucinations, ensures contextual relevance, and mitigates the computational costs inherent in leveraging an LLM.<br>Our results show that Docusage aligns closely with journalist-generated summaries, outperforming foundational and specialized models. Additionally, Docusage offers an interpretable framework that is not constrained by context size, ensures transparency regarding the role of extracted sentences within the narrative, and adapts to the style of the training data.

提供机构：

figshare

创建时间：

2024-10-31

5,000+

优质数据集

54 个

任务类型

进入经典数据集