KdConv
收藏资源简介:
KdConv是一个中文多领域知识驱动对话数据集,由清华大学计算机科学与技术系创建。该数据集包含86,000条语句和4,500个对话,覆盖电影、音乐和旅游三个领域,平均每个对话有19轮。数据集中的对话深入探讨相关话题,并自然过渡到多个话题。创建过程中,通过众包方式收集对话,并要求标注者根据知识图谱中的知识三元组生成语句。KdConv旨在解决多轮知识驱动对话建模的问题,支持知识规划、知识基础和知识适应等研究,适用于探索知识在多轮人机对话中的交互作用。
KdConv is a Chinese multi-domain knowledge-driven dialogue dataset created by the Department of Computer Science and Technology, Tsinghua University. It contains 86,000 utterances and 4,500 dialogues, covering three domains: movies, music and tourism, with an average of 19 turns per dialogue. Dialogues in the dataset thoroughly explore relevant topics and naturally transition across multiple topics. During its creation, dialogues were collected via crowdsourcing, and annotators were required to generate utterances based on knowledge triples in the knowledge graph. KdConv aims to address the problem of multi-turn knowledge-driven dialogue modeling, supports researches including knowledge planning, knowledge grounding and knowledge adaptation, and is applicable to exploring the interactive role of knowledge in multi-turn human-machine dialogues.




