wangyueyiiiiiii/ContextDialog
收藏资源简介:
ContextDialog是一个全面的基准测试,旨在评估语音交互模型在多轮对话中参与、保留和利用相关信息的能力,反映现实世界中人们经常忘记和重新访问过去交流的场景。该数据集基于MultiDialog构建,包含约340小时的数据,涉及12位发言者,每段对话至少10轮。数据集包含测试集,分为test_freq和test_rare两个子集,分别包含363和290段对话,以及1,452和1,160个问答对。数据集的字段包括文件ID、位置、查询类型、分割类型、问题音频、回答音频、问题文本、回答文本和支持文本等。
ContextDialog is a comprehensive benchmark designed to evaluate a voice interaction model’s ability to engage in, retain, and leverage relevant information throughout multi-turn conversations, reflecting real-world scenarios where people often forget and revisit past exchanges. ContextDialog is constructed using MultiDialog, a spoken dialog corpus featuring conversations between two speakers, comprising approximately 340 hours of data with at least 10 turns per conversation from 12 speakers. The dataset includes a test set divided into test_freq and test_rare subsets, containing 363 and 290 dialogues, and 1,452 and 1,160 QA pairs, respectively. The datasets fields include file_id, position, query, split, question_audio, answer_audio, question_text, answer_text, and supporting_text.





