SEACrowd/cod
收藏资源简介:
跨语言基于大纲的对话(COD)数据集由手动生成、本地化和跨语言对齐的任务导向对话(TOD)数据组成,用于对话提示的生成。该数据集支持自然语言理解、对话状态跟踪以及端到端对话建模和评估。数据集由Majewska等人(2022)通过一种新颖的基于大纲的注释管道创建,该管道将英语模式引导对话(SGD)数据集自动采样并映射到大纲中,然后由人类主体进行改写和本地化适配。数据集支持的语言为印尼语(ind),主要任务为对话系统。
The Cross-lingual Outline-based Dialogue (COD) dataset comprises manually generated, localized, and cross-lingually aligned task-oriented dialogue (TOD) data tailored for dialogue prompt generation. This dataset supports natural language understanding, dialogue state tracking, end-to-end dialogue modeling and corresponding evaluation. It was developed by Majewska et al. (2022) via a novel outline-based annotation pipeline, where the English Schema-Guided Dialogue (SGD) dataset is first automatically sampled and mapped into outlines, then revised and localized by human annotators. The dataset supports Indonesian (ind) as its target language, with its core application task focused on dialogue systems.
数据集概述
名称
- Cod
语言
- 印尼语(ind)
任务类别
- 对话系统(Dialogue System)
描述
- Cross-lingual Outline-based Dialogue (COD) 是一个包含手动生成、本地化和跨语言对齐的任务导向对话(TOD)数据集。该数据集支持自然语言理解、对话状态跟踪和端到端对话建模及评估。Majewska 等人(2022)使用一种新颖的基于大纲的注释流程创建了该数据集。
支持的任务
- 对话系统
数据集版本
- 源版本:1.0.0
- SEACrowd 版本:2024.06.20
数据集许可证
- 未知
引用
-
如果使用 Cod 数据集,请引用以下文献:
@article{majewska2022cross, title={Cross-lingual dialogue dataset creation via outline-based generation}, author={Majewska, Olga and Razumovskaia, Evgeniia and Ponti, Edoardo Maria and Vuli{c}, Ivan and Korhonen, Anna}, journal={arXiv preprint arXiv:2201.13405}, year={2022} }




