CDCONV
收藏资源简介:
CDCONV是一个专为中文对话中的矛盾检测设计的基准数据集,由清华大学智能技术与系统国家重点实验室创建。该数据集包含12,000个多轮对话,标注了三种典型的矛盾类别:句子内矛盾、角色混淆和历史矛盾。数据集的创建过程结合了自动对话生成和精细的人工质量筛选,模拟了触发聊天机器人产生矛盾的常见用户行为。CDCONV旨在解决开放域对话系统中的关键问题,即对话矛盾,通过提供丰富的矛盾案例,推动对话模型在理解和处理对话矛盾方面的研究进展。
CDCONV is a benchmark dataset specifically designed for contradiction detection in Chinese conversations, developed by the State Key Laboratory of Intelligent Technology and Systems at Tsinghua University. The dataset comprises 12,000 multi-turn dialogues annotated with three typical contradiction categories: intra-sentence contradiction, role confusion, and historical contradiction. Its construction combines automated dialogue generation and rigorous manual quality screening, simulating common user behaviors that trigger chatbots to generate contradictory outputs. CDCONV aims to address the critical issue of dialogue contradictions in open-domain dialogue systems. By providing a rich collection of contradictory dialogue cases, it advances research on dialogue models' understanding and handling of dialogue contradictions.




