遇见数据集

Syn-D-CNN Dataset

收藏
Figshare2025-02-20 更新2026-04-08 收录
官方服务:

资源简介:

Text summarization condenses extensive content into concise summaries; however, current approaches often rely on large language models (LLMs), which can lack interpretability and are susceptible to generating hallucinated content. To address these issues, we propose Docusage, an interpretable framework that replicates human summaries through a hierarchical clustering approach combined with extractive summarization, augmented by selective, LLM-based abstraction. Docusage minimizes the risk of hallucinations, ensures contextual relevance, and mitigates the computational costs inherent in leveraging an LLM.<br>Our results show that Docusage aligns closely with journalist-generated summaries, outperforming foundational and specialized models. Additionally, Docusage offers an interpretable framework that is not constrained by context size, ensures transparency regarding the role of extracted sentences within the narrative, and adapts to the style of the training data.

提供机构:
Sadmanee, Akib
创建时间:
2024-10-31
二维码
社区交流群
二维码
科研交流群
商业服务