PingPong Benchmark
收藏资源简介:
PingPong Benchmark是由独立研究员Ilya Gusev创建的一个用于评估语言模型角色扮演能力的数据集。该数据集包含288条对话记录,涵盖了多种角色和情境,旨在通过多轮对话模拟真实用户行为并自动评估对话质量。数据集的创建过程结合了系统提示和用户提示,确保了角色的一致性和对话的自然流畅。该数据集主要应用于娱乐领域的语言模型评估,旨在解决模型在互动场景中的角色扮演能力问题。
Created by independent researcher Ilya Gusev, the PingPong Benchmark is a dataset designed to evaluate the role-playing capabilities of language models. It includes 288 dialogue records spanning diverse roles and scenarios, with the goal of simulating real user behaviors via multi-turn conversations and automatically evaluating dialogue quality. The dataset was developed by combining system prompts and user prompts to guarantee consistent character portrayal and natural, fluent dialogue. It is primarily applied to language model evaluation in the entertainment sector, aiming to solve the problem of assessing a model's role-playing performance in interactive scenarios.

- 1PingPong: A Benchmark for Role-Playing Language Models with User Emulation and Multi-Model Evaluation独立研究员 / 阿姆斯特丹 · 2024年



