CARMEM
收藏资源简介:
CARMEM数据集是由宝马集团研究与技术、奥格斯堡大学和慕尼黑工业大学联合创建的合成数据集,专为车载语音助手场景设计。该数据集包含1000个提取对话、1000个检索话语和3000个维护话语,总计5000条数据,平均每个对话包含5.08轮对话和80.78个单词。数据集的生成基于GPT-4模型,确保了对话的多样性和真实性。数据集的主要应用领域是评估车载语音助手的长期记忆系统,旨在解决用户偏好提取、存储和检索的问题,提升个性化用户体验。
The CARMEM dataset is a synthetic dataset jointly developed by BMW Group Research and Technology, the University of Augsburg, and Technical University of Munich, tailored specifically for in-vehicle voice assistant scenarios. It comprises 1000 extraction dialogues, 1000 retrieval utterances, and 3000 maintenance utterances, totaling 5000 samples. On average, each dialogue contains 5.08 conversational turns and 80.78 words. The dataset is generated using the GPT-4 model, which guarantees the diversity and authenticity of the dialogues. Its primary application is to evaluate the long-term memory systems of in-vehicle voice assistants, aiming to address challenges in user preference extraction, storage and retrieval, and ultimately improve personalized user experience.

- 1CarMem: Enhancing Long-Term Memory in LLM Voice Assistants through Category-Bounding宝马集团研究与技术, 奥格斯堡大学, 慕尼黑工业大学 · 2025年



