SHARE
收藏资源简介:
SHARE数据集是由西江大学创建的一个开放领域长期对话数据集,基于电影剧本构建。该数据集不仅包含对话中明确揭示的个人角色信息和事件,还隐含提取了共享记忆。数据集大小为119,087条对话,涵盖了多种电影类型,如浪漫、喜剧和动作。创建过程中,使用了电影剧本解析器和大型语言模型(LLMs)来提取和映射信息。SHARE数据集的应用领域主要在于增强长期对话的连贯性和吸引力,旨在解决现有对话系统在处理长期共享记忆方面的不足。
The SHARE dataset is an open-domain long-term dialogue dataset developed by Xijiang University, built upon film scripts. It not only includes explicit personal character information and events revealed in the dialogues, but also implicitly extracts shared memories. The dataset contains 119,087 dialogues spanning multiple film genres such as romance, comedy, and action. During its development, film script parsers and large language models (LLMs) were employed to extract and map relevant information. The primary application of the SHARE dataset is to enhance the coherence and appeal of long-term dialogues, with the goal of addressing the limitations of current dialogue systems in processing long-term shared memories.
数据集概述
基本信息
- 数据集名称: SHARE
- 数据集类型: JSON
- 数据集地址: https://anonymous.4open.science/r/SHARE-AA1E/SHARE.json
描述
该数据集是一个JSON格式的文件,存储在匿名GitHub平台上。数据集的具体内容未在提供的HTML文本中详细描述。




