"IRWOZ 2.0 - A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations"
收藏资源简介:
"IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domain-specific annotations. However, its initial version contains substantial noise in dialogue states and utterances, limiting state-tracking accuracy. We introduce IRWOZ 2.0, which addresses these limitations through large language model (LLM) enhanced generation (Mistral\/Claude-3.5) and quality refinements. Our improveddataset expands to 390 dialogues across 4 industrial domains (Assembly, Delivery, Position, Relocation), featuring manual corrections and automated typo removal. Benchmark experiments on dialogue state tracking demonstrate significant improvements, with GPT-2\u2019s BLEU-4 score increasing from 0.1651 to 0.5604 compared to original IRWOZ."
IRWOZ通过领域专属标注,实现了工业人机交互(Human-Robot Interaction, HRI)对话系统的优化升级。然而其初始版本在对话状态与话语文本中存在大量噪声,制约了对话状态跟踪的精度。为此我们推出了IRWOZ 2.0,通过大语言模型(Large Language Model, LLM)增强生成(采用Mistral/Claude-3.5模型)与质量优化来解决上述局限。优化后的数据集规模扩展至390段对话,覆盖装配(Assembly)、配送(Delivery)、定位(Position)、搬迁(Relocation)4个工业领域,并涵盖手动校正与自动化错别字修正流程。在对话状态跟踪任务上开展的基准测试结果显示,模型性能得到显著提升:相较于原始IRWOZ,GPT-2的BLEU-4评分从0.1651提升至0.5604。




