vadrami/dialogsum
收藏资源简介:
DIALOGSum Corpus是一个大规模对话摘要数据集,包含13,460个对话(外加100个用于主题生成的保留数据),每个对话都配有相应的人工标注摘要和主题。该数据集旨在支持对话摘要任务,数据来源于多个公开对话语料库,如Dailydialog、DREAM和MuTual,以及一个英语口语练习网站,覆盖了日常生活中的广泛主题,包括学校教育、工作、医疗、购物、休闲和旅行等。对话通常发生在朋友、同事或服务提供者与客户之间,具有清晰的沟通模式和意图,适合用于自动摘要研究。数据集被划分为训练集(12,460个对话)、验证集(500个对话)和测试集(1,500个对话),并包含一个保留集(100个对话,仅含ID、对话和主题字段)。每个数据实例包括对话文本、人工编写的摘要、主题以及唯一ID。数据集的创建基于真实场景,强调传递最突出的信息、简洁性、保留重要命名实体、从观察者视角编写并使用正式语言。
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 dialogues (plus 100 holdout data for topic generation) with corresponding manually labeled summaries and topics. It is designed for summarization tasks and derived from multiple public dialogue corpora, including Dailydialog, DREAM, and MuTual, as well as an English speaking practice website. The dialogues cover a wide range of daily-life topics such as schooling, work, medication, shopping, leisure, and travel, typically occurring between friends, colleagues, or service providers and customers, featuring clear communication patterns and intents suitable for automatic summarization. The dataset is split into train (12,460 dialogues), validation (500 dialogues), test (1,500 dialogues), and a holdout set (100 dialogues with only ID, dialogue, and topic fields). Each instance includes the dialogue text, human-written summary, topic, and a unique ID. The curation rationale emphasizes conveying salient information, brevity, preserving important named entities, writing from an observer perspective, and using formal language.




