MEMORYCODE
收藏资源简介:
MEMORYCODE是一个合成的多会话对话历史数据集,旨在测试大型语言模型在长期交互中跟踪和执行简单编码指令的能力。数据集由多轮对话组成,其中对话者(导师和被指导者)讨论与编码相关的规则和任务。该数据集的特点是包含大量的干扰信息,这些信息与任务无关,但嵌入在对话中,模拟现实工作中的环境。MEMORYCODE通过挑战模型在长对话中检索和更新相关信息的能力,来评估模型在多会话交互中的表现。
MEMORYCODE is a synthetic multi-session dialogue history dataset intended to test the capacity of large language models to track and execute simple coding instructions during long-term interactive scenarios. The dataset comprises multi-turn dialogues where two participants, a tutor and a mentee, discuss coding-related rules and tasks. A prominent feature of this dataset is the incorporation of extensive irrelevant distracting information embedded within the dialogues, which simulates real-world work environments. MEMORYCODE assesses models' performance in multi-session interactions by challenging their ability to retrieve and update relevant information across lengthy dialogues.

- 1From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions巴塞罗那自治大学, 阿姆斯特丹大学, Cohere, Cohere For AI, Cohere For AI Community · 2025年



