Salesforce/RealUserSim
收藏资源简介:
RealUserSim是一个用于现实LLM驱动用户模拟的行为用户档案和评估基准数据集,旨在通过基于真实用户行为的模拟来弥合代理评估中的现实差距。该数据集从WildChat数据集中提取,包含7,273个行为用户档案,每个档案包括人口统计信息(如年龄、性别、教育程度、职业、位置等)和可执行的语言风格命令,用于指导LLM模拟特定用户的沟通风格。此外,数据集提供600个评估测试案例(分为6个领域分割,每个分割100个案例),用于测量用户模拟的保真度,并附带一个LLM-as-a-judge提示,用于在5个行为维度(如人物与情感特征、语言风格与机制、技术能力与知识、交互与数据流、节奏与行动序列)上评估模拟的真实性。基于档案的模拟在600个对话和5个维度上实现了45.3%的保真度,比基线(24.2%)高出21.1个百分点。数据集结构包括档案文件夹(包含用户档案文件)和评估文件夹(包含提示和测试集)。
RealUserSim is a dataset of behavioral user profiles and an evaluation benchmark for realistic LLM-powered user simulation, designed to bridge the reality gap in agent benchmarking via grounded user simulation. Derived from the WildChat dataset, it contains 7,273 behavioral user profiles extracted from real conversations, each including demographics (e.g., age, gender, education, occupation, location) and executable linguistic style commands that guide an LLM to mimic the users communication style. The dataset also provides 600 evaluation test cases (6 splits × 100) for measuring user simulation fidelity, along with an LLM-as-a-judge prompt for evaluating simulation realism across 5 behavioral dimensions, such as Persona & Affective Traits, Linguistic Style & Mechanics, Technical Competency & Knowledge, Interaction & Data Flow, and Pacing & Action Sequencing. Profile-grounded simulation achieves 45.3% fidelity compared to 24.2% for the baseline across 600 conversations and 5 dimensions, representing a 21.1 percentage point improvement. The dataset structure includes a profiles folder (with consolidated user profiles) and an evaluation folder (with the judge prompt and test sets).



