MSCOCO_NAMLab_pt
收藏资源简介:
MultiWOZ_2.2_Chinese是MultiWOZ 2.2数据集的中文翻译版本,专为任务导向对话系统的研究而设计,旨在为中文对话研究提供一个高质量的基准。该数据集涵盖多个领域,包括餐厅、酒店、出租车、景点、医院、警察和火车,包含约10,400个对话,每个对话均通过人工翻译从英文原版转换而来,以确保语言自然性和上下文准确性。数据以JSON格式组织,包括训练集、开发集和测试集,每个对话样本包含对话ID、用户目标、对话历史、系统响应、对话状态等字段。它适用于对话状态跟踪、自然语言生成、端到端对话建模等任务,并可用于评估中文对话系统的性能。数据集的创建过程注重翻译质量,并进行了人工校验以维护一致性。
MultiWOZ_2.2_Chinese is a Chinese translation version of the MultiWOZ 2.2 dataset, specifically designed for research on task-oriented dialogue systems. It aims to provide a high-quality benchmark for Chinese dialogue research, covering multiple domains such as restaurants, hotels, taxis, attractions, hospitals, police, and trains. The dataset contains approximately 10,400 dialogues, each manually translated from the English original to ensure linguistic naturalness and contextual accuracy. Data is organized in JSON format, including training, development, and test sets, with each dialogue sample containing fields such as dialogue ID, user goal, dialogue history, system response, and dialogue state. It is suitable for tasks like dialogue state tracking, natural language generation, end-to-end dialogue modeling, and can be used to evaluate the performance of Chinese dialogue systems. The creation process emphasizes translation quality and has undergone manual verification to maintain consistency.
数据集名称
MSCOCO_NAMLab_pt
许可证
MIT
数据集详情页地址
https://huggingface.co/datasets/Uncertainty-42/MSCOCO_NAMLab_pt
说明
该数据集在Hugging Face平台上托管,由用户Uncertainty-42上传。当前页面未提供关于数据集内容、用途、规模或格式的详细描述。基于名称推测,该数据集可能与Microsoft COCO(Common Objects in Context)数据集相关,并可能经过NAMLab实验室的特定处理或转换(后缀“pt”可能指代PyTorch格式或其他自定义格式),但缺乏官方说明加以确认。




