ConvoSumm
收藏资源简介:
ConvoSumm是一个包含四个子域的对话摘要数据集,由耶鲁大学创建。数据集包括新闻评论、讨论论坛、社区问答和电子邮件线程,每个子域包含250个开发和250个测试示例,总计1000个示例。数据集旨在通过提供多样化的对话形式来推动对话摘要的研究。创建过程中,使用了基于问题-观点-主张框架的注释协议,通过众包方式收集数据。ConvoSumm的应用领域包括自动文本摘要,特别是在理解非结构化对话内容方面,旨在解决如何从多参与者、多视角的对话中提取关键信息的问题。
ConvoSumm is a dialogue summarization dataset with four sub-domains, developed by Yale University. The four sub-domains are news comments, discussion forums, community question answering, and email threads. Each sub-domain includes 250 development set examples and 250 test set examples, resulting in a total of 1000 examples across the entire dataset. This dataset is intended to advance research on dialogue summarization by providing diverse dialogue formats. During its creation, an annotation protocol based on the question-opinion-claim framework was employed, and data was collected via crowdsourcing. The application domains of ConvoSumm include automatic text summarization, particularly for understanding unstructured dialogue content, with the goal of addressing the problem of extracting key information from multi-participant, multi-perspective conversations.




