RHELM (Realistic, Heterogeneous, and Evolving Long-term Memory)
收藏资源简介:
RHELM是由微软与中国人民大学联合构建的面向长期记忆评估的综合性基准数据集,旨在模拟真实世界中的动态用户交互与异构数据流。该数据集包含10个独立人物轨迹,涵盖11,764轮对话和2,180个外部文件,总对话令牌数达477万,外部文件令牌数约243万,数据通过精心设计的用户画像和LOOP模块生成,确保长期语义一致性与时序演化。其核心应用在于评估大型语言模型在复杂现实场景中的记忆能力,特别是解决多源信息聚合、真实上下文推理及隐性状态冲突等关键挑战,推动个性化AI助手的发展。
RHELM is a comprehensive benchmark dataset for long-term memory evaluation, jointly developed by Microsoft and Renmin University of China. It aims to simulate dynamic user interactions and heterogeneous data streams in real-world scenarios. This dataset contains 10 independent character trajectories, covering 11,764 dialogue rounds and 2,180 external files. The total number of dialogue tokens reaches 4.77 million, while the token count of external files is approximately 2.43 million. The data is generated via carefully designed user profiles and the LOOP module, ensuring long-term semantic consistency and temporal evolution. Its core application lies in evaluating the memory capabilities of large language models (LLMs) in complex real-world scenarios, particularly addressing key challenges such as multi-source information aggregation, real-world context reasoning and implicit state conflicts, so as to promote the development of personalized AI assistants.
- 1Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory中国人民大学·应用统计科学研究中心; 中国人民大学·统计学院; 微软 · 2026年



