TwinVoice
收藏资源简介:
TwinVoice是一个用于评估大型语言模型(LLM)在模拟个人身份方面的能力的全面基准数据集。它包含三个维度:社交身份(公共社交互动)、人际身份(私人对话)和叙事身份(基于角色的表达)。该数据集进一步将LLM性能评估分解为六个基本能力,包括观点一致性、记忆回忆、逻辑推理、词汇忠实度、身份语调和句法风格。实验结果表明,尽管先进的模型在身份模拟方面取得了适度的准确度,但它们仍然缺乏句法风格和记忆回忆等能力。因此,LLMs的平均性能仍然远低于人类。该数据集旨在解决当前评估LLM身份模拟能力的局限性,并提供一个系统性的评估框架。
TwinVoice is a comprehensive benchmark dataset for evaluating the ability of Large Language Models (LLMs) to simulate personal identities. It encompasses three dimensions: social identity (public social interactions), interpersonal identity (private conversations), and narrative identity (role-based expression). This dataset further breaks down LLM performance evaluation into six core capabilities, including opinion consistency, memory recall, logical reasoning, lexical fidelity, identity tone, and syntactic style. Experimental results show that although state-of-the-art models have achieved moderate accuracy in identity simulation, they still lack capabilities such as syntactic style and memory recall. Consequently, the average performance of LLMs remains far below that of humans. This dataset aims to address the limitations of current evaluations of LLMs' identity simulation capabilities and provide a systematic assessment framework.
TwinVoice 数据集概述
数据集名称
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
作者信息
- Bangde Du¹* (清华大学)
- Minghao Guo²* (罗格斯大学)
- Songming He³ (复旦大学)
- Ziyi Ye³† (复旦大学)
- Xi Zhu² (罗格斯大学)
- Weihang Su¹ (清华大学)
- Shuqi Zhu¹ (清华大学)
- Yujia Zhou¹ (清华大学)
- Yongfeng Zhang² (罗格斯大学)
- Qingyao Ai¹† (清华大学)
- Yiqun Liu¹ (清华大学)
*同等贡献作者 †通讯作者
项目链接
https://twinvoice.github.io
研究背景
大型语言模型(LLMs)展现出类似人类的涌现能力,被越来越多地设想为模拟个体沟通风格、行为倾向和人格特质的基础。当前基于LLM的人物模拟评估存在局限性:大多依赖合成对话、缺乏系统框架、缺乏能力需求分析。
基准介绍
TwinVoice是一个全面的基准测试,用于评估不同现实场景下的人物模拟能力。该基准涵盖三个维度:
- 社会角色:公共社交互动
- 人际角色:私人对话
- 叙事角色:基于角色的表达
评估能力维度
将LLM性能评估分解为六个基本能力:
- 观点一致性
- 记忆回忆
- 逻辑推理
- 词汇保真度
- 人物语调
- 句法风格
实验发现
实验结果表明,虽然先进模型达到了中等准确率,但在维持一致的人物模拟方面仍然不足,特别是在句法风格和记忆回忆能力方面存在明显欠缺。
任务定义
评估框架:
- LLMs被提示使用特定人物的历史记录并执行刺激任务
- 三种评估协议:
- 判别式:模型从A-D中选择,其中一个是真实人物行为
- 生成式排序:模型生成内容,LLM作为评判者选择最佳候选
- 生成式评分:模型生成内容,评判者在观点、逻辑和风格上评分相似性




