CharacterEval
收藏资源简介:
CharacterEval是一个专为评估中文角色扮演对话代理(RPCA)而设计的大型数据集,由中国人民大学和北京邮电大学的人工智能学院共同创建。该数据集包含1,785个多轮角色扮演对话,总计11,376个示例,涵盖77个来自中国小说和剧本的角色。数据集的构建过程包括使用GPT-4提取对话,随后进行严格的人工质量控制,并通过百度百科补充深入的角色资料。CharacterEval不仅用于评估RPCA的对话能力,还涉及角色一致性、角色扮演吸引力和个性回测等多个维度,旨在全面评估RPCA的性能,解决现有评估方法的不足。
CharacterEval is a large-scale dataset specifically designed for evaluating Chinese Role-Playing Dialogue Agents (RPCA), co-created by the School of Artificial Intelligence of Renmin University of China and Beijing University of Posts and Telecommunications. This dataset contains 1,785 multi-turn role-playing dialogues, totaling 11,376 examples, covering 77 characters sourced from Chinese novels and screenplays. The dataset's construction process includes extracting dialogues using GPT-4, followed by strict manual quality control, and supplementing in-depth character information via Baidu Encyclopedia. CharacterEval is not only used to evaluate the dialogue capabilities of RPCA, but also covers multiple dimensions such as character consistency, role-playing attractiveness, and personality backtesting, aiming to comprehensively evaluate the performance of RPCA and address the shortcomings of existing evaluation methods.

- 1CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation高瓴人工智能学院,中国人民大学 · 2024年



