MIRAGE
收藏资源简介:
MIRAGE数据集由复旦大学大数据研究院和小红书公司联合创建,旨在评估大型语言模型在复杂社交互动环境中的表现。该数据集包含8个独特的剧本,每个剧本具有不同的主题和风格,提供了多样化的模拟环境。数据集的内容包括详细的背景故事和复杂的人际关系网络,支持每个角色的沉浸式角色扮演。数据集的应用领域主要集中在评估模型在复杂社交场景中的表现,旨在解决模型在模拟人类高级行为时的挑战。
The MIRAGE dataset was co-developed by the Institute of Big Data at Fudan University and Xiaohongshu, aiming to evaluate the performance of large language models (LLMs) in complex social interaction environments. This dataset comprises 8 distinct scripts, each with unique themes and styles, offering a diverse range of simulated environments. The dataset features detailed background narratives and intricate interpersonal relationship networks, enabling immersive role-playing for every character. Its primary application scope focuses on evaluating model performance in complex social scenarios, with the objective of addressing the challenges that models face when simulating advanced human behaviors.
MIRAGE 数据集概述
数据集简介
MIRAGE(Multiverse Interactive Role-play Ability General Evaluation)是一个用于评估大型语言模型(LLMs)在复杂角色扮演游戏(如谋杀悬疑游戏)中行为表现的模拟环境。该数据集提供了8个不同的剧本和4种评估方法,用于测试LLMs在复杂社交互动环境中的表现。
数据集内容
剧本信息
数据集包含8个剧本,每个剧本具有不同的结构、类型、结局、阶段数、角色数、线索数以及中英文字数。具体信息如下:
| ID | 剧本名称 | 结构 | 类型 | 结局 | 阶段数 | 角色数 | 线索数 | 中文字数 | 英文字数 |
|---|---|---|---|---|---|---|---|---|---|
| 0 | Bride in filial dress | 单一 | 正统 | 封闭 | 1 | 10 | 39 | 45,475 | 27,503 |
| 1 | The Eastern Star cruise ship | 单一 | 正统 | 开放 | 1 | 5 | 42 | 5,619 | 3,039 |
| 2 | Night at the Museum | 单一 | 非正统 | 封闭 | 1 | 6 | 82 | 13,849 | 6,480 |
| 3 | Li Chuan strange talk book | 单一 | 非正统 | 开放 | 1 | 7 | 14 | 79,012 | 45,666 |
| 4 | The final performance of a big star | 多重 | 正统 | 封闭 | 7 | 2 | 17 | 11,288 | 5,794 |
| 5 | Raging Sea of Rest Life | 多重 | 正统 | 开放 | 2 | 6 | 27 | 18,443 | 6,804 |
| 6 | Article 22 School Rules | 多重 | 非正统 | 封闭 | 5 | 7 | 17 | 91,532 | 41,728 |
| 7 | Fox Hotel | 多重 | 非正统 | 开放 | 2 | 7 | 46 | 107,057 | 62,224 |
评估方法
数据集提供了4种评估方法:
- TII(Trust Inclination Index):信任倾向指数,结合了怀疑和信任分数。
- CIC(Clue Investigation Capability):线索调查能力,衡量LLMs在游戏回合中调查线索的能力。
- ICI(Interactivity Capability Index):互动能力指数,评估LLMs的整体互动能力。
- SCI(Script Compliance Index):剧本遵从指数,评估LLMs在角色扮演中的剧本遵从度。
实验结果
数据集提供了多个模型在MIRAGE场景中的表现结果,具体如下:
| 模型 | Victory | TII | CIC | ICI | SCI | Overall |
|---|---|---|---|---|---|---|
| GPT-3.5 | 29.11 | 47.13 | 27.46 | 70.06 | 49.10 | 44.57 |
| GPT-4 | 34.69 | 76.32 | 19.01 | 76.54 | 50.42 | 51.40 |
| GPT-4o | 47.01 | 78.69 | 35.92 | 76.80 | 51.29 | 57.94 |
| Qwen-2-7B | 51.81 | 75.78 | 18.66 | 74.92 | 50.57 | 54.35 |
| GLM-4-9B | 31.89 | 53.85 | 20.07 | 71.60 | 48.13 | 45.11 |
快速开始
-
安装依赖: bash pip install -r requirements.txt
-
在
config.py中添加API URL和API Key。 -
启动模拟: bash bash run.sh
引用
@article{cai2025mirage, title={MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments}, author={Cai Yin, Gu Zhouhong, Du Zhaohan, Ye Zheyu, Cao Shaosheng, Xu Yiqian, Feng Hongwei, Chen Ping}, journal={arXiv preprint arXiv:2501.01652}, year={2025} }




