CoPrUS-MultiWOZ
收藏资源简介:
CoPrUS-MultiWOZ数据集是由奥斯特拜尔技术大学安贝格-魏登分校的研究团队基于MultiWOZ 2.1数据集创建的,旨在通过添加合成通信错误来增强对话系统的现实性。该数据集包含近1900个对话,通过使用大型语言模型(LLM)生成错误和修复语句,涵盖了误解、非理解和模糊相关问题三种类型的通信错误。创建过程包括两步推理和自动质量保证,以确保生成的语句质量。该数据集主要应用于任务导向对话系统,旨在通过数据增强提高对话系统的泛化能力和用户体验。
The CoPrUS-MultiWOZ dataset was constructed by the research team from the Amberg-Weiden Campus of Ostbayerische Technische Hochschule (OTH) Regensburg, based on the original MultiWOZ 2.1 dataset. Its primary objective is to boost the realism of dialogue systems by introducing synthetic communication errors. Comprising nearly 1,900 dialogues, this dataset encompasses three types of communication errors tied to misunderstanding, non-comprehension and ambiguity, with large language models (LLMs) leveraged to generate erroneous utterances and repair statements during dataset construction. The creation workflow involves two-step inference and automated quality assurance to guarantee the quality of the generated content. Primarily intended for task-oriented dialogue systems, this dataset is designed to enhance the generalization capability and user experience of such systems via data augmentation.

- 1CoPrUS: Consistency Preserving Utterance Synthesis towards more realistic benchmark dialogues奥斯特拜尔技术大学安贝格-魏登分校 · 2024年



