RADDLE
收藏资源简介:
RADDLE是一个专为评估和分析面向任务的对话系统而设计的基准数据集,由微软研究院创建。该数据集包含多个领域的对话任务,旨在评估模型在有限训练数据下的泛化能力和对用户输入不同风格、模态或领域的鲁棒性。RADDLE通过包括具有有限训练数据的任务,鼓励模型发展强大的泛化能力,并提供了一个诊断检查清单,以促进对语言变化、语音错误、未见实体和域外话语等方面的详细鲁棒性分析。此外,RADDLE还提供了一个在线平台,用于模型的评估、比较和鲁棒性分析,旨在解决现有模型在鲁棒性评估中表现不佳的问题,为未来的改进提供机会。
RADDLE is a benchmark dataset developed by Microsoft Research for evaluating and analyzing task-oriented dialogue systems. It covers multi-domain dialogue tasks, aiming to evaluate models' generalization capabilities under limited training data and their robustness against diverse styles, modalities, and domains of user inputs. By incorporating tasks with limited training data, RADDLE encourages models to develop strong generalization abilities, and it also provides a diagnostic checklist to facilitate detailed robustness analysis across linguistic variations, speech errors, unseen entities, out-of-domain utterances and other relevant aspects. Additionally, RADDLE offers an online platform for model evaluation, comparison and robustness analysis, which aims to address the poor performance of existing models in robustness evaluation and provide opportunities for future improvements.

- 1RADDLE: An Evaluation Benchmark and Analysis Platform for Robust Task-oriented Dialog Systems微软研究院 · 2020年



