IAAR-Shanghai/LongEvoRoleBench
收藏资源简介:
LongEvoRoleBench 是一个用于长视野、进化感知角色扮演的统一基准。它标准化了8个现有角色对话语料库,采用共同的下一个话语协议:其中4个长对话语料库测试跨剧集角色状态演化,4个短对话语料库在同一评估格式下提供场景内状态跟踪检查。数据集引入了PHASE-Tree(心理学基础层次属性结构化演化树)作为角色状态表示,并提供原始数据和已处理的分割数据。已处理数据包括多种配置文件表示方法(从M1到M6),其中M6对应完整的PHASE-Tree条件。数据集支持训练和评估,包含训练、随机测试和OOD测试分割,并针对短长期数据集采用不同的OOD策略(短期为未见角色,长期为未见时间周期)。
LongEvoRoleBench is a unified benchmark for long-horizon, evolution-aware role-playing. It standardizes 8 existing character-dialogue corpora into a common next-utterance protocol: 4 long-dialogue corpora test cross-episode character-state evolution, and 4 short-dialogue corpora provide within-scene state-tracking checks under the same evaluation format. The dataset introduces the PHASE-Tree (Psychology-grounded Hierarchical Attribute-Structured Evolving Tree) as the character-state representation and provides both raw source corpora and fully processed evaluation splits. The processed data includes multiple profile representation methods (M1 to M6), with M6 corresponding to the full PHASE-Tree conditioning. It supports training and evaluation with train, random test, and OOD test splits, using different OOD strategies for short-term (unseen characters) and long-term (unseen time periods) datasets.




