KGDS (Knowledge-Grounded Discussion Summarization)
收藏资源简介:
KGDS 数据集旨在解决现有对话摘要系统中由于仅依赖于对话信息而导致外部观察者混淆的问题。该数据集包含 100 个高质量的多领域样本,这些样本来自真实的新闻讨论,并通过严格的标注协议进行标注。每个样本包括结构化的共享背景知识、多轮人工讨论和专家标注的评价组件。该数据集的创建旨在支持对大型语言模型在知识驱动讨论摘要任务上的性能进行评估。
The KGDS dataset is developed to address the issue that existing dialogue summarization systems, which solely leverage dialogue information, cause confusion among external observers. This dataset includes 100 high-quality multi-domain samples sourced from real-world news discussions, annotated under strict annotation protocols. Each sample comprises structured shared background knowledge, multi-turn manual discussions, and expert-annotated evaluation components. The construction of this dataset aims to support the evaluation of Large Language Models (LLMs) on knowledge-driven discussion summarization tasks.

- 1What are they talking about? Benchmarking Large Language Models for Knowledge-Grounded Discussion Summarization北京航空航天大学复杂与关键软件环境国家重点实验室, 中国科学院自动化研究所多模态人工智能系统国家重点实验室, 中国科学院大学, 字节跳动, 番禺人工智能实验室 · 2025年



